Language model cost calculator
Estimate the cost of any language model workload and compare options side by side.
Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.
| Model | Tokens | Per request | Monthly | Context |
|---|---|---|---|---|
| Qwen3.8 Max | 0 est. | $0.003 | $3.00 | 0.1% |
| DeepSeek V4 Flashcheapest | 0 est. | $0.00014 | $0.14 | 0.1% |
| Claude Opus 5 | 0 est. | $0.013 | $12.50 | 0.1% |
Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.
Comparing models on one workload
The fair way to compare language models is to hold the workload fixed (same input and output tokens, same volume) and vary only the model. That isolates price from prompt differences and shows the true cost gap, which is often larger than the headline per-token rates suggest once output pricing is included.
Quality vs cost tradeoffs
A cheaper model that requires a longer prompt to produce the same result, or that needs a retry more often, may not be cheaper in practice. The cost difference between a $1 and $15 per million token model shrinks when the cheaper model uses 30% more tokens and retries 10% of requests. Always benchmark on your actual task before committing to a model for cost reasons.
FAQ
Do cheaper models cost more in the end?
Sometimes. A weaker model may need more tokens, retries or a bigger prompt to hit the same quality, erasing its per-token advantage. Compare on your real task.
How do I benchmark model cost vs quality?
Run 50-100 representative inputs through each candidate model, score the outputs against your quality criteria, and divide the total cost by the number of acceptable outputs. Cost per successful output is the real comparison metric.
What is the most cost-effective model for RAG applications?
For RAG, input tokens dominate since you prepend retrieved context to every query. Models with low input prices and strong instruction-following are most cost-effective. Mid-tier models often outperform budget models on retrieval tasks while remaining much cheaper than flagship models.
Does model size always determine cost?
No. Pricing depends on the provider, the hosting infrastructure, and commercial factors. Some mid-size models are priced higher than larger ones due to specialized capabilities or brand positioning. Always check the actual per-token rate rather than inferring from model size.
Current model pricing
| Model | Provider | Input | Cached | Output | Context | Verified |
|---|---|---|---|---|---|---|
| Qwen3.8 Max | Alibaba | $2 | $0.25 | $6 | 1 000k | 2026-08-03 |
| Claude Haiku 4.5 (latest) | Anthropic | $1 | $0.1 | $5 | 200k | 2025-10-15 |
| Claude Opus 4.8 | Anthropic | $5 | $0.5 | $25 | 1 000k | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic | $3 | $0.3 | $15 | 1 000k | 2026-03-13 |
| DeepSeek Chat | DeepSeek | $0.14 | $0.0028 | $0.28 | 1 000k | 2026-02-28 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.5 | 1 048,576k | 2025-06-17 | |
| Gemini 3.5 Flash | $1.5 | $0.15 | $9 | 1 048,576k | 2026-05-19 | |
| Mistral Medium (latest) | Mistral | $1.5 | n/a | $7.5 | 262,144k | 2026-04-29 |
| Kimi K3 | Moonshot AI | $3 | $0.3 | $15 | 1 048,576k | 2026-07-16 |
| GPT-4.1 mini | OpenAI | $0.4 | $0.1 | $1.6 | 1 047,576k | 2025-04-14 |
| GPT-4o | OpenAI | $2.5 | $1.25 | $10 | 128k | 2024-08-06 |
| GPT-5.5 | OpenAI | $5 | $0.5 | $30 | 1 050k | 2026-04-23 |
| Grok 4.5 | xAI | $2 | $0.3 | $6 | 500k | 2026-07-08 |
| GLM-5.2 | Zhipu AI | $1.4 | $0.26 | $4.4 | 1 000k | 2026-06-13 |
Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.