Token usage calculator
See how many tokens your typical request uses and what that adds up to across a month.
Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.
| Model | Tokens | Per request | Monthly | Context |
|---|---|---|---|---|
| Qwen3.8 Max | 0 est. | $0.003 | $3.00 | 0.1% |
| DeepSeek V4 Flashcheapest | 0 est. | $0.00014 | $0.14 | 0.1% |
| Claude Opus 5 | 0 est. | $0.013 | $12.50 | 0.1% |
Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.
Track usage before it surprises you
Token usage is the metric behind every AI invoice line. Measuring the tokens in a representative request, then multiplying by volume, turns an opaque bill into a number you can plan around and watch for drift as prompts and context grow over time.
How token usage grows over time
Prompts tend to grow as applications mature: system prompts expand with new instructions, conversation history accumulates, and retrieved context gets larger. A prompt that starts at 500 tokens can easily reach 2,000 tokens six months later without anyone noticing. Monitoring token usage per request over time is one of the most important cost-control habits for production AI applications.
FAQ
Why did my token usage go up?
Usually growing conversation history, larger retrieved context, or longer outputs. Re-measure a typical request here and compare against your baseline.
How do I monitor token usage in production?
Most provider APIs return token usage in the response object. Log input_tokens and output_tokens per request, aggregate daily, and alert when the rolling average exceeds your baseline by more than 20%.
What is a typical token usage per API call?
It varies enormously by use case. A simple classification call might use 200 input and 10 output tokens. A document summarization call might use 10,000 input and 500 output tokens. A multi-turn chatbot turn might use 3,000 input (including history) and 300 output tokens.
Can I limit token usage per request?
Yes. All major providers let you set max_tokens (or max_output_tokens) to cap output length. This prevents runaway generation from inflating costs. Set it to the maximum output length you actually need, not the model's hard limit.
Current model pricing
| Model | Provider | Input | Cached | Output | Context | Verified |
|---|---|---|---|---|---|---|
| Qwen3.8 Max | Alibaba | $2 | $0.25 | $6 | 1 000k | 2026-08-03 |
| Claude Haiku 4.5 (latest) | Anthropic | $1 | $0.1 | $5 | 200k | 2025-10-15 |
| Claude Opus 4.8 | Anthropic | $5 | $0.5 | $25 | 1 000k | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic | $3 | $0.3 | $15 | 1 000k | 2026-03-13 |
| DeepSeek Chat | DeepSeek | $0.14 | $0.0028 | $0.28 | 1 000k | 2026-02-28 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.5 | 1 048,576k | 2025-06-17 | |
| Gemini 3.5 Flash | $1.5 | $0.15 | $9 | 1 048,576k | 2026-05-19 | |
| Mistral Medium (latest) | Mistral | $1.5 | n/a | $7.5 | 262,144k | 2026-04-29 |
| Kimi K3 | Moonshot AI | $3 | $0.3 | $15 | 1 048,576k | 2026-07-16 |
| GPT-4.1 mini | OpenAI | $0.4 | $0.1 | $1.6 | 1 047,576k | 2025-04-14 |
| GPT-4o | OpenAI | $2.5 | $1.25 | $10 | 128k | 2024-08-06 |
| GPT-5.5 | OpenAI | $5 | $0.5 | $30 | 1 050k | 2026-04-23 |
| Grok 4.5 | xAI | $2 | $0.3 | $6 | 500k | 2026-07-08 |
| GLM-5.2 | Zhipu AI | $1.4 | $0.26 | $4.4 | 1 000k | 2026-06-13 |
Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.