What the context window is

The context window is the maximum number of tokens a model can consider in a single request. It must hold your system prompt, conversation history, any retrieved context, the user message, and the model’s own reply. Current windows range from a few thousand tokens on older models to a million or more on the largest ones.

What max output is

Max output (sometimes max_tokens or max_output_tokens) is a separate cap on how many tokens the model will generate in one response. It is usually far smaller than the context window. A model with a 128K context might cap output at 4K or 16K tokens.

They share one budget

Input and output draw from the same context window. If the window is 128K and you expect a 4K answer, your prompt must stay under about 124K. Sending a prompt that fills the window leaves no room for the response, so the model truncates or the request fails. Plan the prompt and the answer together.

How to avoid truncation

Keep total input plus expected output under the context window, and set an explicit output cap to the longest answer you actually need. The context window calculator shows your headroom on each model, and the token limit calculator shows what remains after your prompt and expected output.