Context, tokens, and cost
The one unit of measurement that explains pricing, limits, and most weird behaviour.
Models do not read characters or words, they read tokens, chunks roughly three quarters of a word. Everything meters in tokens: what you are charged, how much the model can consider at once, and how long a response takes.
The context window is the total amount the model can have in front of it at one time, counting both your input and its output. It is a hard ceiling, not a guideline, and everything the model knows about the current situation must fit inside it.
- System instructions
Sent on every single call. Small savings here repeat forever. - Retrieved documents
Usually the largest and most controllable part. - Conversation history
Grows without bound unless you manage it. - The output
Typically priced higher than input.
The lever most people miss is caching. Sending the same large preamble on every request often qualifies for a reduced rate, and the identical question asked twice does not need to be answered twice. Both are ordinary engineering, and both are the cheapest wins available.
What to remember
- Tokens are the unit of price, limit, and latency.
- The context window is a hard ceiling shared by input and output.
- Estimate cost per call times realistic volume before shipping.
Terms in this lesson
Field notes
Loaded from a deliberately slow source. The lesson above was already readable while this was still travelling. That is streaming, and it is the same trick a chat interface uses.
The citation that did not exist
A support bot answered a policy question with a confident reference to section 4.2 of a document that has no section 4.2. Both the answer and the citation were generated the same way, which is exactly why one looked as trustworthy as the other.
Grounding matters
A two-cent feature and a four-figure bill
A summarisation feature cost about two cents per use. It was invisible in testing. At launch volume it was several thousand dollars a month, discovered on an invoice rather than in a design review.
Cost per call times realistic volume
resolved in 901ms · region iad1
Hide field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.