LLM pricing calculator

Look up per-token pricing and immediately calculate what your usage costs, no spreadsheets.

0tokens (est.)
0words
0characters
0chars, no spaces

Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.

Models
Qwen3.8 MaxDeepSeek V4 FlashClaude Opus 5
ModelTokensPer requestMonthlyContext
Qwen3.8 Max0 est.$0.003$3.000.1%
DeepSeek V4 Flashcheapest0 est.$0.00014$0.140.1%
Claude Opus 50 est.$0.013$12.500.1%

Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.

Reading per-token pricing

Providers quote prices per million tokens, sometimes per thousand. Input, cached input and output each have their own rate. To compare fairly, normalize everyone to the same basis (this tool uses per 1M) and weight by your own input and output mix rather than trusting a single blended number.

Why the blended rate misleads you

A blended per-token price hides the input/output split. If your workload is 80% output tokens (like a document generation app), a model with cheap input but expensive output will cost far more than its blended rate suggests. Always calculate cost separately for input and output tokens against your actual usage ratio.

FAQ

Is pricing per 1K or per 1M tokens?

This tool normalizes every model to price per 1 million tokens so they are directly comparable. Divide by 1,000 for the per-1K figure.

Why do some models show no cached price?

Not all models support prompt caching. Where no cached rate is shown, the provider either does not offer caching or has not published a separate cached rate. The calculator notes when caching is unavailable for a selected model.

How often is the pricing data updated?

Each model entry has a verifiedAt date linking to the official pricing page. Prices are refreshed when providers announce changes. Always verify against the source link before finalizing a budget.

What is the cheapest model per million tokens?

Budget models like Llama or smaller Mistral variants can be under $0.10 per million input tokens via hosted APIs. The pricing table above sorts cheapest first so you can quickly find the most cost-effective option for your use case.

Current model pricing

ModelProviderInputCachedOutputContextVerified
Qwen3.8 Max Alibaba $2 $0.25 $6 1 000k 2026-08-03
Claude Haiku 4.5 (latest) Anthropic $1 $0.1 $5 200k 2025-10-15
Claude Opus 4.8 Anthropic $5 $0.5 $25 1 000k 2026-05-28
Claude Sonnet 4.6 Anthropic $3 $0.3 $15 1 000k 2026-03-13
DeepSeek Chat DeepSeek $0.14 $0.0028 $0.28 1 000k 2026-02-28
Gemini 2.5 Flash Google $0.3 $0.03 $2.5 1 048,576k 2025-06-17
Gemini 3.5 Flash Google $1.5 $0.15 $9 1 048,576k 2026-05-19
Mistral Medium (latest) Mistral $1.5 n/a $7.5 262,144k 2026-04-29
Kimi K3 Moonshot AI $3 $0.3 $15 1 048,576k 2026-07-16
GPT-4.1 mini OpenAI $0.4 $0.1 $1.6 1 047,576k 2025-04-14
GPT-4o OpenAI $2.5 $1.25 $10 128k 2024-08-06
GPT-5.5 OpenAI $5 $0.5 $30 1 050k 2026-04-23
Grok 4.5 xAI $2 $0.3 $6 500k 2026-07-08
GLM-5.2 Zhipu AI $1.4 $0.26 $4.4 1 000k 2026-06-13

Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.