AI API cost calculator
A focused calculator for AI API spend. Input, output, cached input and monthly requests in, dollars out.
Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.
| Model | Tokens | Per request | Monthly | Context |
|---|---|---|---|---|
| Qwen3.8 Max | 0 est. | $0.003 | $3.00 | 0.1% |
| DeepSeek V4 Flashcheapest | 0 est. | $0.00014 | $0.14 | 0.1% |
| Claude Opus 5 | 0 est. | $0.013 | $12.50 | 0.1% |
Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.
The four inputs that set your bill
An AI API bill comes down to input tokens, output tokens, cached fraction, and requests per month. Get rough numbers for each from a sample request, and the projection is usually within a small margin of your actual invoice, close enough to choose a model and set a budget.
How prompt caching changes your projection
If your system prompt is large and stable, enabling prompt caching can cut input costs by 50-90% for cached tokens. To project this correctly, estimate what fraction of your input tokens are cacheable (typically the system prompt), and apply the cached rate to that fraction. This calculator handles that math automatically when you set the cached percentage.
FAQ
What if I do not know my token counts yet?
Paste a representative prompt into the counter to get exact input tokens, estimate output from a typical response, then set your monthly volume.
How accurate is the monthly projection?
Within a few percent for stable workloads. The main source of variance is output length — if the model generates more tokens than your estimate, costs will be proportionally higher. Build in a 20% buffer for production budgets.
Should I include VAT or taxes in my budget?
Providers typically show pre-tax prices. Your actual invoice may include VAT or local taxes depending on your billing country. Check your provider's invoicing settings and add the applicable tax rate to your projected cost.
Can I use this calculator for multiple models in one workflow?
Yes. Add each model you use in your pipeline as a separate entry and sum the costs. A common pattern is a fast cheap model for classification or routing, plus a powerful model for the actual task — each has its own token volume and cost.
Current model pricing
| Model | Provider | Input | Cached | Output | Context | Verified |
|---|---|---|---|---|---|---|
| Qwen3.8 Max | Alibaba | $2 | $0.25 | $6 | 1 000k | 2026-08-03 |
| Claude Haiku 4.5 (latest) | Anthropic | $1 | $0.1 | $5 | 200k | 2025-10-15 |
| Claude Opus 4.8 | Anthropic | $5 | $0.5 | $25 | 1 000k | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic | $3 | $0.3 | $15 | 1 000k | 2026-03-13 |
| DeepSeek Chat | DeepSeek | $0.14 | $0.0028 | $0.28 | 1 000k | 2026-02-28 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.5 | 1 048,576k | 2025-06-17 | |
| Gemini 3.5 Flash | $1.5 | $0.15 | $9 | 1 048,576k | 2026-05-19 | |
| Mistral Medium (latest) | Mistral | $1.5 | n/a | $7.5 | 262,144k | 2026-04-29 |
| Kimi K3 | Moonshot AI | $3 | $0.3 | $15 | 1 048,576k | 2026-07-16 |
| GPT-4.1 mini | OpenAI | $0.4 | $0.1 | $1.6 | 1 047,576k | 2025-04-14 |
| GPT-4o | OpenAI | $2.5 | $1.25 | $10 | 128k | 2024-08-06 |
| GPT-5.5 | OpenAI | $5 | $0.5 | $30 | 1 050k | 2026-04-23 |
| Grok 4.5 | xAI | $2 | $0.3 | $6 | 500k | 2026-07-08 |
| GLM-5.2 | Zhipu AI | $1.4 | $0.26 | $4.4 | 1 000k | 2026-06-13 |
Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.