Token budget calculator
Set your input and expected output tokens to see the total and how much context headroom is left per model.
Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.
| Model | Tokens | Per request | Monthly | Context |
|---|---|---|---|---|
| Qwen3.8 Max | 0 est. | $0.003 | $3.00 | 0.1% |
| DeepSeek V4 Flashcheapest | 0 est. | $0.00014 | $0.14 | 0.1% |
| Claude Opus 5 | 0 est. | $0.013 | $12.50 | 0.1% |
Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.
Budgeting within a context window
The context window is shared by input and output: prompt, history, retrieved context and the answer all draw from the same budget. Reserve enough room for the response, or the model truncates your context or refuses the request. This tool shows the headroom left on each model after your input and expected output.
Allocating tokens across components
A practical token budget for a typical RAG application might look like: 500 tokens for the system prompt, 200 tokens for the user message, 3,000 tokens for retrieved context chunks, and 500 tokens reserved for output. That totals 4,200 tokens, leaving substantial headroom even on an 8K context model. Map your own components to get a realistic picture before choosing a model.
FAQ
What happens if I exceed the context window?
The provider either rejects the request or silently drops the oldest tokens, which quietly degrades answers. Keep total input plus expected output under the window.
Is a bigger context window always better?
Not necessarily. Larger context windows cost more per request because you are billed for every token you send, including empty space in long contexts. Use a window size that fits your typical workload with some headroom, not the largest available.
How much context should I reserve for output?
Reserve at least the maximum response length you expect. For short answers, 256-512 tokens is usually enough. For detailed explanations or long-form content, reserve 1,000-4,000 tokens. If the model hits the output cap mid-sentence, the response is truncated.
Can I use the full context window for input?
Technically yes, but leaving no room for output will result in truncated responses. Always subtract your expected max output length from the context window to get your usable input budget.
Current model pricing
| Model | Provider | Input | Cached | Output | Context | Verified |
|---|---|---|---|---|---|---|
| Qwen3.8 Max | Alibaba | $2 | $0.25 | $6 | 1 000k | 2026-08-03 |
| Claude Haiku 4.5 (latest) | Anthropic | $1 | $0.1 | $5 | 200k | 2025-10-15 |
| Claude Opus 4.8 | Anthropic | $5 | $0.5 | $25 | 1 000k | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic | $3 | $0.3 | $15 | 1 000k | 2026-03-13 |
| DeepSeek Chat | DeepSeek | $0.14 | $0.0028 | $0.28 | 1 000k | 2026-02-28 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.5 | 1 048,576k | 2025-06-17 | |
| Gemini 3.5 Flash | $1.5 | $0.15 | $9 | 1 048,576k | 2026-05-19 | |
| Mistral Medium (latest) | Mistral | $1.5 | n/a | $7.5 | 262,144k | 2026-04-29 |
| Kimi K3 | Moonshot AI | $3 | $0.3 | $15 | 1 048,576k | 2026-07-16 |
| GPT-4.1 mini | OpenAI | $0.4 | $0.1 | $1.6 | 1 047,576k | 2025-04-14 |
| GPT-4o | OpenAI | $2.5 | $1.25 | $10 | 128k | 2024-08-06 |
| GPT-5.5 | OpenAI | $5 | $0.5 | $30 | 1 050k | 2026-04-23 |
| Grok 4.5 | xAI | $2 | $0.3 | $6 | 500k | 2026-07-08 |
| GLM-5.2 | Zhipu AI | $1.4 | $0.26 | $4.4 | 1 000k | 2026-06-13 |
Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.