LLM pricing calculator
Look up per-token pricing and immediately calculate what your usage costs, no spreadsheets.
Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.
| Model | Tokens | Per request | Monthly | Context |
|---|---|---|---|---|
| Qwen3.8 Max | 0 est. | $0.003 | $3.00 | 0.1% |
| DeepSeek V4 Flashcheapest | 0 est. | $0.00014 | $0.14 | 0.1% |
| Claude Opus 5 | 0 est. | $0.013 | $12.50 | 0.1% |
Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.
Reading per-token pricing
Providers quote prices per million tokens, sometimes per thousand. Input, cached input and output each have their own rate. To compare fairly, normalize everyone to the same basis (this tool uses per 1M) and weight by your own input and output mix rather than trusting a single blended number.
Why the blended rate misleads you
A blended per-token price hides the input/output split. If your workload is 80% output tokens (like a document generation app), a model with cheap input but expensive output will cost far more than its blended rate suggests. Always calculate cost separately for input and output tokens against your actual usage ratio.
FAQ
Is pricing per 1K or per 1M tokens?
This tool normalizes every model to price per 1 million tokens so they are directly comparable. Divide by 1,000 for the per-1K figure.
Why do some models show no cached price?
Not all models support prompt caching. Where no cached rate is shown, the provider either does not offer caching or has not published a separate cached rate. The calculator notes when caching is unavailable for a selected model.
How often is the pricing data updated?
Each model entry has a verifiedAt date linking to the official pricing page. Prices are refreshed when providers announce changes. Always verify against the source link before finalizing a budget.
What is the cheapest model per million tokens?
Budget models like Llama or smaller Mistral variants can be under $0.10 per million input tokens via hosted APIs. The pricing table above sorts cheapest first so you can quickly find the most cost-effective option for your use case.
Current model pricing
| Model | Provider | Input | Cached | Output | Context | Verified |
|---|---|---|---|---|---|---|
| Qwen3.8 Max | Alibaba | $2 | $0.25 | $6 | 1 000k | 2026-08-03 |
| Claude Haiku 4.5 (latest) | Anthropic | $1 | $0.1 | $5 | 200k | 2025-10-15 |
| Claude Opus 4.8 | Anthropic | $5 | $0.5 | $25 | 1 000k | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic | $3 | $0.3 | $15 | 1 000k | 2026-03-13 |
| DeepSeek Chat | DeepSeek | $0.14 | $0.0028 | $0.28 | 1 000k | 2026-02-28 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.5 | 1 048,576k | 2025-06-17 | |
| Gemini 3.5 Flash | $1.5 | $0.15 | $9 | 1 048,576k | 2026-05-19 | |
| Mistral Medium (latest) | Mistral | $1.5 | n/a | $7.5 | 262,144k | 2026-04-29 |
| Kimi K3 | Moonshot AI | $3 | $0.3 | $15 | 1 048,576k | 2026-07-16 |
| GPT-4.1 mini | OpenAI | $0.4 | $0.1 | $1.6 | 1 047,576k | 2025-04-14 |
| GPT-4o | OpenAI | $2.5 | $1.25 | $10 | 128k | 2024-08-06 |
| GPT-5.5 | OpenAI | $5 | $0.5 | $30 | 1 050k | 2026-04-23 |
| Grok 4.5 | xAI | $2 | $0.3 | $6 | 500k | 2026-07-08 |
| GLM-5.2 | Zhipu AI | $1.4 | $0.26 | $4.4 | 1 000k | 2026-06-13 |
Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.