Context window
AI in the stackHow much a model can consider at once.
The total amount of text (input and output together) a model can have in front of it for a single call. A hard limit. When a conversation exceeds it, earlier messages must be dropped or summarised.
See also Token (model), Prompt caching