Token usage calculator

See how many tokens your typical request uses and what that adds up to across a month.

0tokens (est.)
0words
0characters
0chars, no spaces

Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.

Models
Qwen3.8 MaxDeepSeek V4 FlashClaude Opus 5
ModelTokensPer requestMonthlyContext
Qwen3.8 Max0 est.$0.003$3.000.1%
DeepSeek V4 Flashcheapest0 est.$0.00014$0.140.1%
Claude Opus 50 est.$0.013$12.500.1%

Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.

Track usage before it surprises you

Token usage is the metric behind every AI invoice line. Measuring the tokens in a representative request, then multiplying by volume, turns an opaque bill into a number you can plan around and watch for drift as prompts and context grow over time.

How token usage grows over time

Prompts tend to grow as applications mature: system prompts expand with new instructions, conversation history accumulates, and retrieved context gets larger. A prompt that starts at 500 tokens can easily reach 2,000 tokens six months later without anyone noticing. Monitoring token usage per request over time is one of the most important cost-control habits for production AI applications.

FAQ

Why did my token usage go up?

Usually growing conversation history, larger retrieved context, or longer outputs. Re-measure a typical request here and compare against your baseline.

How do I monitor token usage in production?

Most provider APIs return token usage in the response object. Log input_tokens and output_tokens per request, aggregate daily, and alert when the rolling average exceeds your baseline by more than 20%.

What is a typical token usage per API call?

It varies enormously by use case. A simple classification call might use 200 input and 10 output tokens. A document summarization call might use 10,000 input and 500 output tokens. A multi-turn chatbot turn might use 3,000 input (including history) and 300 output tokens.

Can I limit token usage per request?

Yes. All major providers let you set max_tokens (or max_output_tokens) to cap output length. This prevents runaway generation from inflating costs. Set it to the maximum output length you actually need, not the model's hard limit.

Current model pricing

ModelProviderInputCachedOutputContextVerified
Qwen3.8 Max Alibaba $2 $0.25 $6 1 000k 2026-08-03
Claude Haiku 4.5 (latest) Anthropic $1 $0.1 $5 200k 2025-10-15
Claude Opus 4.8 Anthropic $5 $0.5 $25 1 000k 2026-05-28
Claude Sonnet 4.6 Anthropic $3 $0.3 $15 1 000k 2026-03-13
DeepSeek Chat DeepSeek $0.14 $0.0028 $0.28 1 000k 2026-02-28
Gemini 2.5 Flash Google $0.3 $0.03 $2.5 1 048,576k 2025-06-17
Gemini 3.5 Flash Google $1.5 $0.15 $9 1 048,576k 2026-05-19
Mistral Medium (latest) Mistral $1.5 n/a $7.5 262,144k 2026-04-29
Kimi K3 Moonshot AI $3 $0.3 $15 1 048,576k 2026-07-16
GPT-4.1 mini OpenAI $0.4 $0.1 $1.6 1 047,576k 2025-04-14
GPT-4o OpenAI $2.5 $1.25 $10 128k 2024-08-06
GPT-5.5 OpenAI $5 $0.5 $30 1 050k 2026-04-23
Grok 4.5 xAI $2 $0.3 $6 500k 2026-07-08
GLM-5.2 Zhipu AI $1.4 $0.26 $4.4 1 000k 2026-06-13

Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.