AI in the stackLesson 2 of 48 min

Context, tokens, and cost

The one unit of measurement that explains pricing, limits, and most weird behaviour.

Models do not read characters or words, they read tokens, chunks roughly three quarters of a word. Everything meters in tokens: what you are charged, how much the model can consider at once, and how long a response takes.

The context window is the total amount the model can have in front of it at one time, counting both your input and its output. It is a hard ceiling, not a guideline, and everything the model knows about the current situation must fit inside it.

Where the tokens go
  1. System instructions
    Sent on every single call. Small savings here repeat forever.
  2. Retrieved documents
    Usually the largest and most controllable part.
  3. Conversation history
    Grows without bound unless you manage it.
  4. The output
    Typically priced higher than input.

The lever most people miss is caching. Sending the same large preamble on every request often qualifies for a reduced rate, and the identical question asked twice does not need to be answered twice. Both are ordinary engineering, and both are the cheapest wins available.

What to remember

  • Tokens are the unit of price, limit, and latency.
  • The context window is a hard ceiling shared by input and output.
  • Estimate cost per call times realistic volume before shipping.

Terms in this lesson

Show field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.