AI cost calculator

Model your real AI bill. Set expected input, output, cached percentage and requests per month, then compare providers.

0tokens (est.)
0words
0characters
0chars, no spaces

Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.

Models
Qwen3.8 MaxDeepSeek V4 FlashClaude Opus 5
ModelTokensPer requestMonthlyContext
Qwen3.8 Max0 est.$0.003$3.000.1%
DeepSeek V4 Flashcheapest0 est.$0.00014$0.140.1%
Claude Opus 50 est.$0.013$12.500.1%

Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.

What drives your AI bill

Total spend is cost per request multiplied by request volume. The levers are prompt size, output length, how much input you can cache, and how many calls you make per month. Small per-request savings compound fast at scale, so the model that looks cheapest per token is not always cheapest for your workload once volume is factored in.

Hidden costs to watch for

Beyond token prices, AI costs include retry calls (failed requests still bill partial tokens in some cases), evaluation runs during development, and agent steps that make multiple LLM calls per user action. A chatbot that runs 3 LLM calls per user message triples your effective token spend versus a single-call setup.

FAQ

How do I lower my AI costs?

Trim the prompt, cache stable prefixes, cap output length, and route routine tasks to a cheaper model. Compare the same workload across models above to find the best fit.

Which AI model is cheapest overall?

For simple tasks, smaller open-weight models like Llama or Mistral hosted on budget providers can cost a fraction of flagship models. For complex reasoning, a powerful model that solves the task in one attempt may be cheaper than a cheap model that requires multiple retries.

How do I estimate costs before I build?

Write a representative prompt, paste it into the token counter to get input token count, estimate your typical output length, then multiply by your expected monthly request volume. That gives a realistic pre-launch cost estimate.

Do AI costs scale linearly with users?

Roughly yes, if each user generates similar request volumes. Costs scale with total tokens processed, not with user count directly. A power user making 100 requests costs 100x more than a user making 1 request, regardless of seat count.

Current model pricing

ModelProviderInputCachedOutputContextVerified
Qwen3.8 Max Alibaba $2 $0.25 $6 1 000k 2026-08-03
Claude Haiku 4.5 (latest) Anthropic $1 $0.1 $5 200k 2025-10-15
Claude Opus 4.8 Anthropic $5 $0.5 $25 1 000k 2026-05-28
Claude Sonnet 4.6 Anthropic $3 $0.3 $15 1 000k 2026-03-13
DeepSeek Chat DeepSeek $0.14 $0.0028 $0.28 1 000k 2026-02-28
Gemini 2.5 Flash Google $0.3 $0.03 $2.5 1 048,576k 2025-06-17
Gemini 3.5 Flash Google $1.5 $0.15 $9 1 048,576k 2026-05-19
Mistral Medium (latest) Mistral $1.5 n/a $7.5 262,144k 2026-04-29
Kimi K3 Moonshot AI $3 $0.3 $15 1 048,576k 2026-07-16
GPT-4.1 mini OpenAI $0.4 $0.1 $1.6 1 047,576k 2025-04-14
GPT-4o OpenAI $2.5 $1.25 $10 128k 2024-08-06
GPT-5.5 OpenAI $5 $0.5 $30 1 050k 2026-04-23
Grok 4.5 xAI $2 $0.3 $6 500k 2026-07-08
GLM-5.2 Zhipu AI $1.4 $0.26 $4.4 1 000k 2026-06-13

Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.