Token budget calculator

Set your input and expected output tokens to see the total and how much context headroom is left per model.

0tokens (est.)
0words
0characters
0chars, no spaces

Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.

Models
Qwen3.8 MaxDeepSeek V4 FlashClaude Opus 5
ModelTokensPer requestMonthlyContext
Qwen3.8 Max0 est.$0.003$3.000.1%
DeepSeek V4 Flashcheapest0 est.$0.00014$0.140.1%
Claude Opus 50 est.$0.013$12.500.1%

Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.

Budgeting within a context window

The context window is shared by input and output: prompt, history, retrieved context and the answer all draw from the same budget. Reserve enough room for the response, or the model truncates your context or refuses the request. This tool shows the headroom left on each model after your input and expected output.

Allocating tokens across components

A practical token budget for a typical RAG application might look like: 500 tokens for the system prompt, 200 tokens for the user message, 3,000 tokens for retrieved context chunks, and 500 tokens reserved for output. That totals 4,200 tokens, leaving substantial headroom even on an 8K context model. Map your own components to get a realistic picture before choosing a model.

FAQ

What happens if I exceed the context window?

The provider either rejects the request or silently drops the oldest tokens, which quietly degrades answers. Keep total input plus expected output under the window.

Is a bigger context window always better?

Not necessarily. Larger context windows cost more per request because you are billed for every token you send, including empty space in long contexts. Use a window size that fits your typical workload with some headroom, not the largest available.

How much context should I reserve for output?

Reserve at least the maximum response length you expect. For short answers, 256-512 tokens is usually enough. For detailed explanations or long-form content, reserve 1,000-4,000 tokens. If the model hits the output cap mid-sentence, the response is truncated.

Can I use the full context window for input?

Technically yes, but leaving no room for output will result in truncated responses. Always subtract your expected max output length from the context window to get your usable input budget.

Current model pricing

ModelProviderInputCachedOutputContextVerified
Qwen3.8 Max Alibaba $2 $0.25 $6 1 000k 2026-08-03
Claude Haiku 4.5 (latest) Anthropic $1 $0.1 $5 200k 2025-10-15
Claude Opus 4.8 Anthropic $5 $0.5 $25 1 000k 2026-05-28
Claude Sonnet 4.6 Anthropic $3 $0.3 $15 1 000k 2026-03-13
DeepSeek Chat DeepSeek $0.14 $0.0028 $0.28 1 000k 2026-02-28
Gemini 2.5 Flash Google $0.3 $0.03 $2.5 1 048,576k 2025-06-17
Gemini 3.5 Flash Google $1.5 $0.15 $9 1 048,576k 2026-05-19
Mistral Medium (latest) Mistral $1.5 n/a $7.5 262,144k 2026-04-29
Kimi K3 Moonshot AI $3 $0.3 $15 1 048,576k 2026-07-16
GPT-4.1 mini OpenAI $0.4 $0.1 $1.6 1 047,576k 2025-04-14
GPT-4o OpenAI $2.5 $1.25 $10 128k 2024-08-06
GPT-5.5 OpenAI $5 $0.5 $30 1 050k 2026-04-23
Grok 4.5 xAI $2 $0.3 $6 500k 2026-07-08
GLM-5.2 Zhipu AI $1.4 $0.26 $4.4 1 000k 2026-06-13

Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.