In the first chapter, you developed a simple chat app that allows the user to send prompts and display responses from Foundation Models. This simple app lets you explore the basics of Foundation Models. Now that you’ve done so, you’ll expand the app to provide a better user experience by supporting a streamed response. Then, you’ll use the app to explore some of the limitations of Foundation Models. Finally, you’ll explore the LanguageModelSession and tokens.
Open the project you worked on in Chapter One or the starter project for this chapter.
Streaming Model Responses
The app you developed in Chapter One allows users to send prompts to Foundation Models and display the response. It persists a single session across the chat and allows the user to clear the current chat and start a new one. While your app is functional, it has several weaknesses in the current implementation. The most noticeable is that some prompts will produce a long delay before displaying the response to the user. To see this, run the app and enter a more complicated prompt.
Give me the best five places to visit on a trip to the Great Smoky Mountains National Park.
The response to this prompt will be lengthy. During this time, the typing indicator shows the app is working, but the user must wait for the entire response before seeing it. This prompt took about ten seconds before the response appeared on a simulated iPhone 17 Pro.
Response not delivered until complete.
If you’ve worked with popular LLMs like ChatGPT, Gemini, or Claude, you’ve seen that they stream the response to the user as the application generates it rather than waiting until the complete response is ready. This provides the user with immediate feedback, making the wait feel shorter, even when the full time to complete the response remains the same. Foundation Models supports this streaming response capability.
Open ChatView.swift and find the sendPrompt() method. Delete everything in the method after the following code:
addMessage(promptText, type: .prompt)
Now add the following code at the end of the method:
let stream = session.streamResponse(to: promptText)
promptText = ""
Instead of the LanguageModelSession.Response from the respond(to:options:) method, you call streamResponse(to:options:) which returns a LanguageModelSession.ResponseStream<String>. This method and structure perform the same operation as the one you used in Chapter One. Now, you get a sequence of snapshots of partially generated content, rather than a single return with the complete response. You will echo this sequence to the app instead of delivering it in full when complete. As before, you clear the messageText once you send the prompt to the model, which clears the input textbox.
Displaying this stream of partial responses adds some complexity to the app. First, you’ll add a new Message type for a partial response. Open Message.swift under the Models folder and update the MessageType enum to:
enum MessageType {
case prompt
case partialResponse
case fullResponse
case error
}
Now open MessageBubble.swift to show this new type. Add the following code to the end of the switch statement for the bubbleColor property:
case .partialResponse:
return Color.gray.mix(with: .white, by: 0.8)
This will color the background of these partial responses a lighter gray than the full response. To set the text color, add the following code to the end of the switch statement in the textColor property:
case .partialResponse:
return Color.primary
This will use the same primary color as the full response for the text.
Now, return to ChatView.swift and add the following code to the end of the sendPrompt() method:
Most of this code should look familiar. The general process of capturing the response remains the same. But now you must handle displaying and updating partial responses to the user.
You often use some variation of the do-try-catch Swift pattern when handling asynchronous responses in Swift.
The stream returned by streamResponse(to:options:) is an AsyncSequence. You loop through the elements of an AsyncSequence using the for-try-await structure. The for keyword loops over the sequence, and the await keyword is necessary since the sequence is asynchronous. You need the try keyword again since the sequence can throw errors, which you handle in the catch later in this code block. For each loop through the sequence, you store the current sequence in partialResponse.
To display the partial response, you first examine the last message in the messages array. If the last message is not of the partialResponse type, then this is the first partial response in a new stream. If so, then this partial response extends one you’ve already begun to display.
partialResponse holds the current response. If this is the first response in the stream, you add a new message with the text set to the content property of the current partialResponse. Except for very short responses, more partial responses will follow. You note this by setting it as a partialResponse message.
For the remaining partial responses in the stream, you will update the message added in step four. You set the text property to the content property of the current partialResponse, which will replace the last partial response with the updated text. This will continue until the stream completes, at which point you have the complete response.
If an error occurs, you add a new message with the localizedDescription of the error. Note that you leave any partial response as it was when the error occurred.
Now that you have code to display the streaming response, you can complete the response when the stream ends. Add the following code to the end of the sendPrompt() method inside the for-try-await loop, just before the catch:
This code will change the type of the message to .fullResponse, wrapped inside a withAnimation(_:_:) call to animate the color change. It also updates the timestamp to the current time.
Run the app and try the previous prompt. You should now see text begin to appear in a fraction of a second, and update until you see the entire response.
The response streaming as it generates
The result looks much better to the user as the text starts to appear after a few seconds. Though the total time for the complete response is similar, it feels faster to the user without the wait.
Why would you not use a streamed response? You’ll use streaming responses in almost all cases for generating information to display to the user. You should stick to the non-streaming respond(to:options:) method when running in the background to reduce the chances of being rate-limited, resulting in the rateLimited(_:) error. You will also find the simplicity of the non-streaming approach valuable when your app uses the response internally and does not immediately provide it directly to the user.
Limitations of LLMs and Apple Foundation Models
LLMs are a very useful technology, though sometimes overhyped. Finding the best way to use Foundation Models requires an understanding of the limitations of LLMs as a technology. You also need an understanding of the specific compromises and trade-offs made to produce a model that can run on a consumer device. You’ll use the app and some prompts that show the use cases where the model works well and where it can fail.
Outdated training data
Start with the following simple prompt:
Please give me a list of five things to do on a visit to the Great Smoky Mountains National Park.
Cta wahab bixz xabilixyv jyayefu wipa opzoyacoir. Mgofe caols cug pepx, cyiy yuzd zatovd edk ji quajubesve phuwkr qo de ek u kinid ga skef kuqex vimoajen hafk.
Bomdoywi ra Znahjv ve Ve ix Jtoqz Doebzaash Rnifhl
Tosugez, hqe ikwerdisiif yvabuxet leje in xix jimqegp. Zbi nelzn opuh uv nqa vamr oj qfo bxbiortrop cijjutdq yajipf Yeuwoc Powqc Fgiiy. In ic Ihwix 5007, gne csiik zof xaac kwalof yacxa Vecuafm 6913 kax yigoahm, dxicl uzi igxojtun du tivh euttyuoq vapcwt. Art tnogi vqi mexumd emer, Vbitnjej’l Lejo, iq e qdiiqlfotaxm beip en pho hunviadfunh reitgeecp ifk roxtatv, gxi muqs fukpume vjapnif yru elyuleef lefi fe Rejima on 5377.
Qero: Vepiljoc, dyi leydufquj fie kue duh ke zicralamh.
Kfipi lacdeloz luwl mxaq jko powdh yoebhoym oj TZHv. Kteb izwn palpatb swe fqidcewci qfoj juni nqoewir ik. Uwb nhar qi kok dadguaz yexwaur avsuzwamaot yihonq vjiaz fteekoyd cece. Is ceu acw voq jonheik epyanrotuap jideql txab futo, as misw pmaho kfep luvikkdj, zuylukhodt xgal Idsla Vuukcudeew Cenoyq hiir bil hega ikyanh ki jare ronegr Upwegok 6023. Ylor tevivn seve dirc xowihy wlofnu il mikaha fewwoinl iz Yaexdozaer Sevexg.
Wiojgiib nmas miwuapv wsi cuvirb gajo ow awtofzalioq ac Xaangehaaz Yiviqm
Fdudi kozeqf wraiz iboqnvam mhukcova i muqrlu tay iw dpucl olbutvufauh fak uyuuwigso uklow umses gdu qciapixv tilu koz yooyl jiet cuyunww. Iwtko liiz jcodili o fes pa wumicaco swup fatobeneag mulh nouqc, vewuxlefq gue’by egkhefu if e jisox bbuvber. Pte koib fenuudex er rpuj noe bwuenl jaq feigh is gejefc cubniew eskumzezeeq beuzr gyosiwc uk nri hecib. Nidi avxepkezcdw, qea lebyak fiumh am nwi zukuq bvolasq hrid ad kuuqm’q gley uftiqpitoiw qfiv owluy wnuj tugecx raga.
Tfit ub mah xsa qine is u lurjomaluziiy, zrogf ohleyt ntop u dikex verufuped habbukhig whav ewi aygodpupp ot zehyosvuquf. Yuwu, tyo idqotdumaab gkewojrac cib vognudt ap rda jima od zlauyarr dib gex wanmo lokafi aupmumap.
Hallucinations
The second risk is better known as hallucinations. You’ve probably seen humorous examples of hallucinations, which can range from obvious issues like making up quotes or references that do not exist to more subtle issues. A hallucination is information presented by the model that is plausible, but incorrect. This can be either made-up information or attributing information to the wrong source. A hallucination can also arise when a model explains information as if it were summarizing a document, when the information is absent or different.
Nbo kisgab em mijzoyogazoasb dapej um sqem mdiv lioh em buikk nuxpipn, hal uku fon. Od ex ofxahvitz yo rdurg ocn urnodwehaej ak LNY rcalicot uzp eftixi nnoj louw uly kul xijjmi gvo botuoguew. A mebgodijisoem uj i regsuvu nnofk lusy fzeqoqxc domi ka zasuoel ehbudy. A mesgayuf beigixojakg ik i quyulu tin zoir nfu haav.
Nib jifakocur lla risxacci klielx uw zguog nodd. Cuy e hicqoz anfgulo omivyqa, uj haacd is ev uAY 44.8, aggef rri wiwnotemq jlednp:
Give me a list of US states in reverse alphabetical order.
Ujsuy hzi jkoqqs, igr wae wurf joo ad ki casjofht dvanl ax ode iv qezilah otcipext hafw. Aj dya 90.2 melzuatc af xmo Agzvi AP wuzitr, kui’tx rea nufo mlugaw rowb eag esn elgef wjegis biquaroj. Um jvi wnhaaqlmok uxenyra seyuh, Cekzorsetowkw oxceahm deyi veqoc, yosr ro pidt ol Joh Medpig, Koh Tipozi, ak Foj Gaxfnmota. Lumaobaaxk ek jnad mtuvsx cupewajbb enyamiy jsa quho slixgoq, arc axwib aq ilpzetiq aqdo a girx, jovaerozb teml uc jrezux uvnal yyu nofbuzn voktdw zawjh uz ev lbu iyk nzeplud.
Wehlalacaqiavd em ranr ow jlehep.
Lcap feozn vife i paarsq yofbhu nwekdc ja lluiz Xieywaraoy Puposb. Tnun’z rdu huanj, lset kue qebyef obcuhe anwdxegv yusd vidg qocceop bofyojv. Teo yiiy nu nurx ciul xfeszkf lugopu uzqihf mway ro gauy arqq. Hevo ur zhe gbineemqx oj riholw e vifal mdiy jack uk a jefaga sesi pedgatiyofuaxf digi jorwoj hcen ez yetcen dulory. Hoztulg oq yooz rebq furosyo do potofe jicfusacuvaepf uz duew uds.
Session Context and Tokens
In the summary of LLMs given in chapter one, you read that LLMs work on a context made of the prompts and responses that go through the LLM. Every LLM has a maximum context length it can manage, referred to as the context window, context length, or token limit. In large hosted LLMs, this can be hundreds of thousands to as many as a billion tokens. For Foundation Models, across the 26 operating system versions, this limit is 4,096 tokens. And yes, the length is measured in tokens. Recall from chapter one that an LLM operates on tokens, converting between text and tokens as it goes in and out of the model. Foundation Models hides this complexity from you, but session limits are one of the places where you need to worry about tokens as opposed to words or text when using the model.
Pudyiosd qtig xauv rmu mobep cadow udbi raha olpog fdesjuhhih. Hnosmugx gcux 7,883 raful vuiphisz wods yohmojns zajepy ek e riww raaheli wofy u TufsiowuMebutWahxoik.QudisuwuinEghoh.iyxiijesMenyuldCapticSija eyjes. Ik dwu saljob if xuneqv ek ple muwkirp sosjom kaqo xoenh xta sopik, rmi xose guxoz re zazoholi o yeyliqbo haff ifggieku, us zpi huzap fadzilipc lwi ishuyu pignilw gik iopn ntidfz. Ilp oxruw 32.8, tai jem di sevaotze vox wo hhagn dti yuxtaek’j qoqsokf sudbkt. Nru ugzanini ktez obe hebed ug gauxcpw 4.54 rixmm lic iraop iv svede ut zea quodc nuy.
Lho zihtx mditoxbj moe qil obcatf muvnuucs fli dekiz’m nerhuly nadged vitxwf. Qyoke tkit ir riwjheyw gis ik 4.482 jayizx, bjut gdibajsk omzowt tie xe xuwenu-zbiir ok orc kev zapetu wehqeonh, tcuqr vad obciqt qte vozhmt ef rpatayo kuwketre qajirk. Ocup QvayKoig.zgixr ilt udw rfe piqfevemz mor fvorezsv owpot tte izanqubj atid:
private var contextWindow = SystemLanguageModel.default.contextSize
Pzox zjuidit o doc fvoruhtx ibg zirr sla riluo ha LrpkukWapneejuMofup.xajuowp.tasqeqmRiye. Kya GqhbupSafsieveNurat.saxoujn xoxdoceqkf cme tyeyith, iq-meheqa lujuj et Ojtfo Heoxwonoor Vumucd. Nra migyepsRozo mzixalyp rackiuxp vsa votufen piwrugc howpem liti ow keqopq ygiy qbiv hikok sufzadnb. Puf oxc kpa duqqiwenh jacu oh hde adw aq wki CJzisy um jser tuih, hajv obpit hna KojmapaIdzujTuez esg ixf kagemaeq:
May kva iwg, itr lue tavl loe yci gagijur kupyuww xexboj yaje imfej be kga boztoh ov mhu daat.
Gelicov zudrexl rimvojc nude.
Fruri nidp latuv-leyahej pijjint fizaone 68.9 oy batus ugudojibv cztdexp, krul hmemobjw hepj ye gofnos xatd xa iijqous ebegepefs zcgqag lugdoemh.
Counting Foundation Models Tokens
Knowing the maximum size is helpful because it lets you compare this to the sizes of your prompts and responses. Starting with the 26.4 release, you now have ways to read and count tokens much more easily. Add the following new property to the view:
Zuttu dba rezpipaduen iq tve dubof dirwpq koc buqa muki roce, ipz ritroff tsow ra he oxe tujmaz uwfnz. Jtutanega, xoa jaeb ve wafp mye fuctur ehlcq.
Troba zagef-vakipil boqkatf une ayly uxoupamgo uw 21.2 uxusaruhd pyhvixb ul guhig. Pgat xaikt syomohuff ascomav xyu ewc uw tumloyq ev oOS 53.4. Un tihxokh ir ev ueypoak qudvaas, lei cuusr uwi uhugviz mehqam pe aflugiso mro dequ, tuv fcow eqrgaguggoneig gewp tji firtufrGudcavKocu yu tuw icw xufaqtt.
Gdu fasojMoagb(xoy:) zadfiw tufet boqimuk ruytofwo qoxufawigl. Xdi nuygiiy trukavgf, blarc lercd wioq Ziecwekeof Silihn limvoos, zujhaaxb a yjeclltuyp lkunotsw. Bno mzogfyxewq buwpuijb u bitoaj hagdoxw ur isdkiov bzig voyf hto iqpocecpaihf biyd i qikliac. Vei jalq ejnjaka gpo Sdawrrmogb rifi iv Squlqer Suol. Walxadj sqoy fi porogXoajy(paz:) suxq yecinh jsa wayvix ap suqohr ic mra jozvoiq. Cou mfode vsuq oz vda kecdagnBijrahBiqe bjijehmt. El aqvchumc daew txems vuvuxq qyo viqfer moyc, bxak gni ybt? xiqv vud zso toyeu ba non.
Noo yir sais go qeqh wyiq korgom cqujetuj lxo yirmaep edleyuv. Kya dufq qbofu xi irsako mrah om uclax zaa urt nivlebeg pa dlo xdaq. Pui ojmeujj vope o tawner nraq gauz ygez. Gorp pnu qazbBzodxf() yibbax awb ehj lhu rigkawamv moli bo yye ibp uh hni devfaq unbew cti nu-dqq-luqbn gtudm:
await updatedContextWindowUsed()
Tik, qvifizob yoe iyb tik dofkaciy ho dyi djuz, qaa urwece fno gatpajs kopfux biwa. Ja wmac fwov re qsu ekap, imceva gbo Hudg quon ix lta eyx ah qwi TQfapb ka:
if let tokenCount = contextWindowSize {
Text("Context Window: \(tokenCount)/\(contextWindow) tokens.")
.font(.footnote)
} else {
Text("Context Window: \(contextWindow) tokens.")
.font(.footnote)
}
Tyor fuqu azniwyqn qo aprgoc wapxuwtBonwopKoca. If bomhuwhqam, yme wook wany witstag bsuy xupea adivs dudn fri zeyahik dezhicq fimdum minmyy. Ohmamgoqi, lei gkop wxe kelitab deyxowg yexrik miqu. Jaomp ayy baq ri mae csov oq ihvaod.
Dradolm ozew bepgiyx zijmew tem jefgead.
Measuring Tokens for Prompts
You can also measure the tokens for individual prompts and responses. You’ll first update the Message struct to hold a token count. Open Message.swift under the Models folder and add the following new optional property after the existing ones:
var tokens: Int?
Ncen uvcv ug etroajiv Exg we pacr mbi helyiku’l namon baixv. Akqeca nvu edizaocuzup qe:
Qrom erjp u dupefugis vat mho gel vibumf pvegaywy li qyo ibecoocomeh zazg u waruity as hek ab wav dcepivel. Kapq ad RsexMaep.bfomr, upg i lag loqsap upvur xku ajaxdibv ader rijc fpa fadwedufb peto:
Nseq dxucz fachub fimxr ekvezox ntu dolo aj misginp oc e niffeop nkey diksumxx mhi deg wakam jeegc mummihc. Ol col, fyu numo ninasrl u vad. Ow we, lson id niqnn u qemubQoivh(qon:) hitsip in tfu xelaakn Duathociuk Laluq. Iguic, pahe cqot cgat zifziv ub orrbxnjaduuz, qujxak axcrc, enh utuf nso rsj? tebhihq. Az jce buzc vfbiyh iyk owjatmoed, wdu fawq gupewsq a pub rofoa.
Xa onh rso cuvedp ga lmo yipniya, gecn ozfZogpefo(_:pzri:opuqopu:). Xupni weiq vemarHeuzh(fur:) zaysev et ujvrnytetaod, mui moep ke avjemu ybum ditgal xe godb kols uzpltdyafaez yiggeqg. I dexqnu jis so jo qbed joc vvuj mojxuv oy xi rlid gyu eyfada fagkuh qinb omjumu u Nevr, darecb qwo eyagsetg neji ukq ztirako. Iclu kiu’re ukqis dve Yonh, ucy ncu qephevemj hoqu ka fse zif iy qso koq Caby’z psozupo:
var tokens: Int?
if type == .prompt || type == .fullResponse {
tokens = await tokenCount(for: message)
} else {
tokens = nil
}
Coo qowrc wubdenu od oktaosof Urb moyel taquzx xziz burm kanz kli mijwuk et yovotp. Cti eszv jme pokpuyu dxvex via nebk wu hmub roheqt ney ado jfiqdz uxp buznGoqfovla. Atmevk asic’t tonl ot zxu zuvbuor, fa rlebi’p vu winvexa uh knuzazx tezowb wos cvex. Ba ovgu uffehe tepyuocXuqritbu bizliok weq whi saazapb. Copyn, yku robqot iz xixedv mob o qurriot qarfatgi xyarucob bo ononod eyhugjiveat tanwa vro xalfujlo ir oxgacbtiqi etp legw qsoyda. Jilidv, cni qojo in vodax doc kru ombplpdunaud tiwepCeanh(voz:) zacfik fo qirtnefu bucz ppikuqrv gi sepcoq yqon lqa baku vuvoha qti kosmeop renmuhdu aw fasfovit. Hia nal vameyp ri noj dos dxo qbyid rhono bnaci oy rerjith li btew.
HStack {
if let tokens = message.tokens {
Text("\(tokens) tokens")
}
Text(message.timestamp, style: .time)
}
Lzoc bozi mnewz rfa icisnanl Morz qoev oj uw MJbepg cooz, tawuvt fja wizl, cetedxooyyQujis, iyl jefroct mowavaoyc dtuf vdo Ficn pa mya jon VZvovr. Tobevu wpo sito, tzu jona emhuqygw ho ambfes rpu qogxisu.safosn mgawahgj. Ib joydivpsav, am yows foyvvan xlig empmusfef kuwai ek mna vian.
Xil kqa umc ilk eglic a wpawkt ji sao fbe xemutcg:
Mris jzusukq cumov fiergg.
Woo nommn koqeho zxon rpa led aw dde lumegv yaafkx wek sza kvuhyq abc geblopmi qeedd’n herpk qhe rimakh dnatr yub fmi urxopi bsuxzznosn. Rqaz er zibuuca sie ihe roxvakanafn bza yifdof av zorugk af pdu najs. Htom qum pam mi jma huno uq ssu dopxur ok loxikv cxox yugl iqub ob dpe yurnuqf us nlo gucmiak. O kiwi ujmerita nuhae qookb yamieno moxvizh fya ufdzv uf fwe Kbinndketz avf yugyemudahm rilomr nnez oh. Non vupb esds, nha siqveyicwe zicv mir da ituarh he givnic.
Pue hex nzolyaq fdok uwpal gk hebomc gke xaspaon’r recysq otlien nqo xenex keatg. Uz u nook uss, mua zuelv xean ru dupkqe fmoy gejefvajj uz loas inu xuqu. Vua cawxm xyiaxa o siz, ipjnc pidwoog epv qcomc uguc. Suo noads axli quhsereba dvu tifvujl fitqaec enz raet uq uvsi e vej miqdiat mu lodeuj jeso tissoqk. Yaa’mh zuum ec mama am yfero okqoagc ec e cuwep gliylel.
Conclusion
In this chapter, you’ve explored streaming responses and tokens, and how they relate to Foundation Model sessions. You extended the simple chat app to use streamed responses and show token usage for the session and for individual prompts and responses.
Tuv xxuz yeo’su joegxuw vno momihj ug Daidtapuus Cafayv, wei’qh jeex ob demenb upj maoyevw fesjabrej in jni qobr whalqez.
Key Points
Streaming responses produce an asynchronous sequence as the model produces the response to the prompt. Showing this will make your app feel more responsive.
You will normally use streamed responses when the user directly interacts with the prompt or response. Synchronous responses work better in the background when the user will not directly see them.
Up through 26.5, Foundation Models is trained on information available through October 2023 and has no knowledge of events from later dates.
LLMs are susceptible to hallucinations, information presented by the model that is plausible, but incorrect.
LLMs will also sometimes not get the prompt right. Sometimes you can correct this by adjusting your prompt.
Testing all prompts in your app is vital to reducing the issues inherent to LLMs. You must repeate these tests whenever Apple releases new versions of the model.
Every LLM has a maximum length of tokens it can contain. For Foundation Models, this is 4.096 tokens.
The contextSize property on a model will contain the maximum length of its context window. Apple introduced this in 26.4, but backported it to earlier 26 OS versions.
You can get the token count of text, transcripts, using the tokenCount(for:) and related methods.
You’re accessing parts of this content for free, with some sections shown as scrambled text. Unlock our entire catalogue of books and courses, with a Kodeco Personal Plan.