The three token types

Your AI bill is built from three separately-priced token types:

Input tokens
Everything you send: system prompt, conversation history, retrieved context and the user’s message. Usually the largest and cheapest per-token portion.
Output tokens
What the model generates. Priced higher than input, often 3–5×, so long responses cost more than long prompts.
Cached input tokens
Input the provider has seen recently and can reuse. Billed at a steep discount (often 10–25% of the normal input price). Worth it when a large, stable prefix (a long system prompt or fixed context) repeats across many requests.

When caching saves money

Set a cached-input percentage in the calculator to see the savings on your own workload. Not every model supports caching, and the calculator flags those. The bigger and more stable your repeated prefix, the more caching helps; a short or constantly-changing prompt gains little.