Token cost calculator

Enter input and output tokens to see the exact cost per request and per month for each model you select.

0tokens (est.)
0words
0characters
0chars, no spaces

Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.

Models
Qwen3.8 MaxDeepSeek V4 FlashClaude Opus 5
ModelTokensPer requestMonthlyContext
Qwen3.8 Max0 est.$0.003$3.000.1%
DeepSeek V4 Flashcheapest0 est.$0.00014$0.140.1%
Claude Opus 50 est.$0.013$12.500.1%

Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.

How token cost is calculated

Cost is billed per token, split into three rates: input (your prompt), cached input (a reused prefix, billed at a steep discount where supported), and output (what the model generates). Output usually costs three to five times more than input, so a short prompt with a long answer can cost more than a long prompt with a short answer.

Input vs output pricing explained

Providers charge more for output tokens because generating text requires more GPU compute than reading it. A model that processes your 500-token prompt and returns a 1,000-token answer is doing most of its work on the output side. This is why summarization tasks (long input, short output) are much cheaper than generation tasks (short input, long output).

FAQ

Why is output more expensive than input?

Generating tokens is more compute-intensive than reading them, so most providers price output several times higher than input. Capping max output tokens is one of the fastest ways to cut cost.

What is cached input pricing?

Some providers offer a discounted rate for input tokens that come from a recently cached prompt prefix. If your system prompt stays the same across requests, the provider can cache it and bill subsequent calls at a lower rate, sometimes 50-90% cheaper than regular input.

How much does 1 million tokens cost?

It depends heavily on the model. Budget models can cost under $0.10 per million input tokens, while flagship reasoning models can cost $15 or more. Output is typically 3-5x the input price. Use the calculator above to compare exact costs for your chosen models.

Does the cost change if I send the same prompt repeatedly?

Yes, if prompt caching is enabled. Providers that support caching will bill the first call at the full input rate and subsequent calls at the cached rate for the unchanged prefix. This can significantly reduce cost for high-volume applications with a fixed system prompt.

Current model pricing

ModelProviderInputCachedOutputContextVerified
Qwen3.8 Max Alibaba $2 $0.25 $6 1 000k 2026-08-03
Claude Haiku 4.5 (latest) Anthropic $1 $0.1 $5 200k 2025-10-15
Claude Opus 4.8 Anthropic $5 $0.5 $25 1 000k 2026-05-28
Claude Sonnet 4.6 Anthropic $3 $0.3 $15 1 000k 2026-03-13
DeepSeek Chat DeepSeek $0.14 $0.0028 $0.28 1 000k 2026-02-28
Gemini 2.5 Flash Google $0.3 $0.03 $2.5 1 048,576k 2025-06-17
Gemini 3.5 Flash Google $1.5 $0.15 $9 1 048,576k 2026-05-19
Mistral Medium (latest) Mistral $1.5 n/a $7.5 262,144k 2026-04-29
Kimi K3 Moonshot AI $3 $0.3 $15 1 048,576k 2026-07-16
GPT-4.1 mini OpenAI $0.4 $0.1 $1.6 1 047,576k 2025-04-14
GPT-4o OpenAI $2.5 $1.25 $10 128k 2024-08-06
GPT-5.5 OpenAI $5 $0.5 $30 1 050k 2026-04-23
Grok 4.5 xAI $2 $0.3 $6 500k 2026-07-08
GLM-5.2 Zhipu AI $1.4 $0.26 $4.4 1 000k 2026-06-13

Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.