AI in the stackLesson 2 of 48 min

Context, tokens, and cost

The one unit of measurement that explains pricing, limits, and most weird behaviour.

Models do not read characters or words, they read tokens, chunks roughly three quarters of a word. Everything meters in tokens: what you are charged, how much the model can consider at once, and how long a response takes.

The context window is the total amount the model can have in front of it at one time, counting both your input and its output. It is a hard ceiling, not a guideline, and everything the model knows about the current situation must fit inside it.

Where the tokens go
  1. System instructions
    Sent on every single call. Small savings here repeat forever.
  2. Retrieved documents
    Usually the largest and most controllable part.
  3. Conversation history
    Grows without bound unless you manage it.
  4. The output
    Typically priced higher than input.

The lever most people miss is caching. Sending the same large preamble on every request often qualifies for a reduced rate, and the identical question asked twice does not need to be answered twice. Both are ordinary engineering, and both are the cheapest wins available.

What to remember

  • Tokens are the unit of price, limit, and latency.
  • The context window is a hard ceiling shared by input and output.
  • Estimate cost per call times realistic volume before shipping.

Terms in this lesson

Field notes

Loaded from a deliberately slow source. The lesson above was already readable while this was still travelling. That is streaming, and it is the same trick a chat interface uses.

The citation that did not exist

A support bot answered a policy question with a confident reference to section 4.2 of a document that has no section 4.2. Both the answer and the citation were generated the same way, which is exactly why one looked as trustworthy as the other.

Grounding matters

A two-cent feature and a four-figure bill

A summarisation feature cost about two cents per use. It was invisible in testing. At launch volume it was several thousand dollars a month, discovered on an invoice rather than in a design review.

Cost per call times realistic volume

resolved in 901ms · region iad1

Hide field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.