AI cost calculator
Model your real AI bill. Set expected input, output, cached percentage and requests per month, then compare providers.
Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.
| Model | Tokens | Per request | Monthly | Context |
|---|---|---|---|---|
| Qwen3.8 Max | 0 est. | $0.003 | $3.00 | 0.1% |
| DeepSeek V4 Flashcheapest | 0 est. | $0.00014 | $0.14 | 0.1% |
| Claude Opus 5 | 0 est. | $0.013 | $12.50 | 0.1% |
Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.
What drives your AI bill
Total spend is cost per request multiplied by request volume. The levers are prompt size, output length, how much input you can cache, and how many calls you make per month. Small per-request savings compound fast at scale, so the model that looks cheapest per token is not always cheapest for your workload once volume is factored in.
Hidden costs to watch for
Beyond token prices, AI costs include retry calls (failed requests still bill partial tokens in some cases), evaluation runs during development, and agent steps that make multiple LLM calls per user action. A chatbot that runs 3 LLM calls per user message triples your effective token spend versus a single-call setup.
FAQ
How do I lower my AI costs?
Trim the prompt, cache stable prefixes, cap output length, and route routine tasks to a cheaper model. Compare the same workload across models above to find the best fit.
Which AI model is cheapest overall?
For simple tasks, smaller open-weight models like Llama or Mistral hosted on budget providers can cost a fraction of flagship models. For complex reasoning, a powerful model that solves the task in one attempt may be cheaper than a cheap model that requires multiple retries.
How do I estimate costs before I build?
Write a representative prompt, paste it into the token counter to get input token count, estimate your typical output length, then multiply by your expected monthly request volume. That gives a realistic pre-launch cost estimate.
Do AI costs scale linearly with users?
Roughly yes, if each user generates similar request volumes. Costs scale with total tokens processed, not with user count directly. A power user making 100 requests costs 100x more than a user making 1 request, regardless of seat count.
Current model pricing
| Model | Provider | Input | Cached | Output | Context | Verified |
|---|---|---|---|---|---|---|
| Qwen3.8 Max | Alibaba | $2 | $0.25 | $6 | 1 000k | 2026-08-03 |
| Claude Haiku 4.5 (latest) | Anthropic | $1 | $0.1 | $5 | 200k | 2025-10-15 |
| Claude Opus 4.8 | Anthropic | $5 | $0.5 | $25 | 1 000k | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic | $3 | $0.3 | $15 | 1 000k | 2026-03-13 |
| DeepSeek Chat | DeepSeek | $0.14 | $0.0028 | $0.28 | 1 000k | 2026-02-28 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.5 | 1 048,576k | 2025-06-17 | |
| Gemini 3.5 Flash | $1.5 | $0.15 | $9 | 1 048,576k | 2026-05-19 | |
| Mistral Medium (latest) | Mistral | $1.5 | n/a | $7.5 | 262,144k | 2026-04-29 |
| Kimi K3 | Moonshot AI | $3 | $0.3 | $15 | 1 048,576k | 2026-07-16 |
| GPT-4.1 mini | OpenAI | $0.4 | $0.1 | $1.6 | 1 047,576k | 2025-04-14 |
| GPT-4o | OpenAI | $2.5 | $1.25 | $10 | 128k | 2024-08-06 |
| GPT-5.5 | OpenAI | $5 | $0.5 | $30 | 1 050k | 2026-04-23 |
| Grok 4.5 | xAI | $2 | $0.3 | $6 | 500k | 2026-07-08 |
| GLM-5.2 | Zhipu AI | $1.4 | $0.26 | $4.4 | 1 000k | 2026-06-13 |
Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.