Prompt token calculator
Paste your prompt to count its tokens exactly (GPT) or as an estimate, and see how much of the context window it uses.
Token count uses the exact GPT (o200k) tokenizer. Counted in your browser, never uploaded.
| Model | Tokens | Per request | Monthly | Context |
|---|---|---|---|---|
| Qwen3.8 Max | 0 est. | $0.003 | $3.00 | 0.1% |
| DeepSeek V4 Flashcheapest | 0 est. | $0.00014 | $0.14 | 0.1% |
| Claude Opus 5 | 0 est. | $0.013 | $12.50 | 0.1% |
Token counts vary by tokenizer and model; non-OpenAI counts are estimates. Cost is an estimate and provider pricing can change. See each model's source on the pricing page.
What counts as prompt tokens
Your prompt is more than the user message: it includes the system prompt, any few-shot examples, tool definitions and the full conversation history resent on every turn. All of it is billed as input on every call, so a large fixed system prompt is often the biggest and most cacheable part of the bill.
Reducing prompt token count
The most effective ways to shrink a prompt are: remove redundant instructions (models follow concise prompts better anyway), replace few-shot examples with fine-tuning or RAG, summarize conversation history rather than resending it in full, and compress tool definitions to only the fields the model needs. A 30% reduction in prompt size translates directly to a 30% reduction in input cost.
FAQ
Does the system prompt count every request?
Yes. The system prompt and any history are resent and billed as input on every call, unless the provider caches a stable prefix at the discounted cached rate.
How many tokens does a typical system prompt use?
Simple system prompts run 50-200 tokens. Detailed prompts with role instructions, output format rules, and examples can reach 500-2,000 tokens. Enterprise prompts with extensive policy rules sometimes exceed 5,000 tokens.
Do tool definitions count as input tokens?
Yes. Every tool or function definition you pass to the API is tokenized and billed as input. If you have 10 tools defined and only use one per call, you are still paying for all 10 to be read on every request. Only pass the tools the model actually needs for each call.
How do I shrink a prompt without losing quality?
Start by removing filler phrases and redundant instructions. Then check if any few-shot examples can be replaced by a clearer instruction. Finally, see if history can be summarized rather than resent in full. Each step can cut 20-40% of tokens with little quality loss.
Current model pricing
| Model | Provider | Input | Cached | Output | Context | Verified |
|---|---|---|---|---|---|---|
| Qwen3.8 Max | Alibaba | $2 | $0.25 | $6 | 1 000k | 2026-08-03 |
| Claude Haiku 4.5 (latest) | Anthropic | $1 | $0.1 | $5 | 200k | 2025-10-15 |
| Claude Opus 4.8 | Anthropic | $5 | $0.5 | $25 | 1 000k | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic | $3 | $0.3 | $15 | 1 000k | 2026-03-13 |
| DeepSeek Chat | DeepSeek | $0.14 | $0.0028 | $0.28 | 1 000k | 2026-02-28 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.5 | 1 048,576k | 2025-06-17 | |
| Gemini 3.5 Flash | $1.5 | $0.15 | $9 | 1 048,576k | 2026-05-19 | |
| Mistral Medium (latest) | Mistral | $1.5 | n/a | $7.5 | 262,144k | 2026-04-29 |
| Kimi K3 | Moonshot AI | $3 | $0.3 | $15 | 1 048,576k | 2026-07-16 |
| GPT-4.1 mini | OpenAI | $0.4 | $0.1 | $1.6 | 1 047,576k | 2025-04-14 |
| GPT-4o | OpenAI | $2.5 | $1.25 | $10 | 128k | 2024-08-06 |
| GPT-5.5 | OpenAI | $5 | $0.5 | $30 | 1 050k | 2026-04-23 |
| Grok 4.5 | xAI | $2 | $0.3 | $6 | 500k | 2026-07-08 |
| GLM-5.2 | Zhipu AI | $1.4 | $0.26 | $4.4 | 1 000k | 2026-06-13 |
Prices per 1M tokens (USD). Each model links to its official source. Seed values pending re-verification. See the methodology.