Context, tokens, and cost
The one unit of measurement that explains pricing, limits, and most weird behaviour.
Models do not read characters or words, they read tokens, chunks roughly three quarters of a word. Everything meters in tokens: what you are charged, how much the model can consider at once, and how long a response takes.
The context window is the total amount the model can have in front of it at one time, counting both your input and its output. It is a hard ceiling, not a guideline, and everything the model knows about the current situation must fit inside it.
- System instructions
Sent on every single call. Small savings here repeat forever. - Retrieved documents
Usually the largest and most controllable part. - Conversation history
Grows without bound unless you manage it. - The output
Typically priced higher than input.
The lever most people miss is caching. Sending the same large preamble on every request often qualifies for a reduced rate, and the identical question asked twice does not need to be answered twice. Both are ordinary engineering, and both are the cheapest wins available.
What to remember
- Tokens are the unit of price, limit, and latency.
- The context window is a hard ceiling shared by input and output.
- Estimate cost per call times realistic volume before shipping.
Terms in this lesson
Show field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.