Why it makes things up
Confident wrong answers are the expected behaviour, not a malfunction, and retrieval is the mitigation.
A language model produces plausible continuations. It has no separate store of facts it consults and no internal signal distinguishing "I know this" from "this is the shape an answer would take". A fabricated citation and a real one are produced by exactly the same process, which is why they look equally convincing.
The standard mitigation is retrieval: before answering, search your own trusted content, put the relevant passages into the prompt, and instruct the model to answer only from them. This is what people mean by RAG, and structurally it is nothing exotic: a search step, then a model call.
- Question arrives
From a user. - Search
Find the handful of passages most relevant to it. - Assemble
Put those passages and the question into one prompt. - Answer with sources
The model responds from the supplied text, and you show which passages it used.
Showing sources matters even more than it appears. It gives the user a way to check, it gives you a way to debug, and it changes the failure mode from an invisible falsehood into a visible mismatch between claim and citation.
What to remember
- Fabrication is inherent to how the model works, not a bug to be patched.
- Fluency and accuracy are unrelated signals.
- Retrieval grounds answers, and its quality caps the whole system.
Terms in this lesson
Show field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.