Token cost calculator
Enter input and output tokens to see the exact cost per request and per month for each model you select.
Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.
| Model | Tokens | Per request | Monthly | Context |
|---|---|---|---|---|
| Qwen3.8 Max | 0 est. | $0.003 | $3.00 | 0.1% |
| DeepSeek V4 Flashcheapest | 0 est. | $0.00014 | $0.14 | 0.1% |
| Claude Opus 5 | 0 est. | $0.013 | $12.50 | 0.1% |
Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.
How token cost is calculated
Cost is billed per token, split into three rates: input (your prompt), cached input (a reused prefix, billed at a steep discount where supported), and output (what the model generates). Output usually costs three to five times more than input, so a short prompt with a long answer can cost more than a long prompt with a short answer.
Input vs output pricing explained
Providers charge more for output tokens because generating text requires more GPU compute than reading it. A model that processes your 500-token prompt and returns a 1,000-token answer is doing most of its work on the output side. This is why summarization tasks (long input, short output) are much cheaper than generation tasks (short input, long output).
FAQ
Why is output more expensive than input?
Generating tokens is more compute-intensive than reading them, so most providers price output several times higher than input. Capping max output tokens is one of the fastest ways to cut cost.
What is cached input pricing?
Some providers offer a discounted rate for input tokens that come from a recently cached prompt prefix. If your system prompt stays the same across requests, the provider can cache it and bill subsequent calls at a lower rate, sometimes 50-90% cheaper than regular input.
How much does 1 million tokens cost?
It depends heavily on the model. Budget models can cost under $0.10 per million input tokens, while flagship reasoning models can cost $15 or more. Output is typically 3-5x the input price. Use the calculator above to compare exact costs for your chosen models.
Does the cost change if I send the same prompt repeatedly?
Yes, if prompt caching is enabled. Providers that support caching will bill the first call at the full input rate and subsequent calls at the cached rate for the unchanged prefix. This can significantly reduce cost for high-volume applications with a fixed system prompt.
Current model pricing
| Model | Provider | Input | Cached | Output | Context | Verified |
|---|---|---|---|---|---|---|
| Qwen3.8 Max | Alibaba | $2 | $0.25 | $6 | 1 000k | 2026-08-03 |
| Claude Haiku 4.5 (latest) | Anthropic | $1 | $0.1 | $5 | 200k | 2025-10-15 |
| Claude Opus 4.8 | Anthropic | $5 | $0.5 | $25 | 1 000k | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic | $3 | $0.3 | $15 | 1 000k | 2026-03-13 |
| DeepSeek Chat | DeepSeek | $0.14 | $0.0028 | $0.28 | 1 000k | 2026-02-28 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.5 | 1 048,576k | 2025-06-17 | |
| Gemini 3.5 Flash | $1.5 | $0.15 | $9 | 1 048,576k | 2026-05-19 | |
| Mistral Medium (latest) | Mistral | $1.5 | n/a | $7.5 | 262,144k | 2026-04-29 |
| Kimi K3 | Moonshot AI | $3 | $0.3 | $15 | 1 048,576k | 2026-07-16 |
| GPT-4.1 mini | OpenAI | $0.4 | $0.1 | $1.6 | 1 047,576k | 2025-04-14 |
| GPT-4o | OpenAI | $2.5 | $1.25 | $10 | 128k | 2024-08-06 |
| GPT-5.5 | OpenAI | $5 | $0.5 | $30 | 1 050k | 2026-04-23 |
| Grok 4.5 | xAI | $2 | $0.3 | $6 | 500k | 2026-07-08 |
| GLM-5.2 | Zhipu AI | $1.4 | $0.26 | $4.4 | 1 000k | 2026-06-13 |
Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.