Prompt caching
AI in the stackReusing an unchanged prefix across calls.
When the same large preamble is sent on every request, providers often charge a reduced rate for the repeated part. One of the cheapest available reductions in model cost.
See also Token (model), Context window