AI in the stackLesson 3 of 48 min

Why it makes things up

Confident wrong answers are the expected behaviour, not a malfunction, and retrieval is the mitigation.

A language model produces plausible continuations. It has no separate store of facts it consults and no internal signal distinguishing "I know this" from "this is the shape an answer would take". A fabricated citation and a real one are produced by exactly the same process, which is why they look equally convincing.

The standard mitigation is retrieval: before answering, search your own trusted content, put the relevant passages into the prompt, and instruct the model to answer only from them. This is what people mean by RAG, and structurally it is nothing exotic: a search step, then a model call.

Retrieval, in order
  1. Question arrives
    From a user.
  2. Search
    Find the handful of passages most relevant to it.
  3. Assemble
    Put those passages and the question into one prompt.
  4. Answer with sources
    The model responds from the supplied text, and you show which passages it used.

Showing sources matters even more than it appears. It gives the user a way to check, it gives you a way to debug, and it changes the failure mode from an invisible falsehood into a visible mismatch between claim and citation.

What to remember

  • Fabrication is inherent to how the model works, not a bug to be patched.
  • Fluency and accuracy are unrelated signals.
  • Retrieval grounds answers, and its quality caps the whole system.

Terms in this lesson

Field notes

Loaded from a deliberately slow source. The lesson above was already readable while this was still travelling. That is streaming, and it is the same trick a chat interface uses.

The citation that did not exist

A support bot answered a policy question with a confident reference to section 4.2 of a document that has no section 4.2. Both the answer and the citation were generated the same way, which is exactly why one looked as trustworthy as the other.

Grounding matters

A two-cent feature and a four-figure bill

A summarisation feature cost about two cents per use. It was invisible in testing. At launch volume it was several thousand dollars a month, discovered on an invoice rather than in a design review.

Cost per call times realistic volume

resolved in 900ms · region iad1

Hide field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.