Per-token pricing

LLM providers bill for usage, not seats. You pay for the tokens you send and the tokens the model generates. Prices are quoted per million tokens (sometimes per thousand). Because every request consumes tokens, your bill scales with total tokens processed, not with how many users you have.

Input, output and cached rates

There are three separately-priced rates. Input is your prompt and any resent context. Output is what the model generates, priced higher because generation is the expensive part. Cached input is a reused prefix billed at a discount where the provider supports it. The input, output and cached tokens guide covers each in detail.

Relative price of input, cached input and output tokens Three bars showing output tokens cost the most, input less, and cached input the least. Cached input Input Output Output is usually the priciest; cached input the cheapest
Relative per-token cost of the three token types. Exact rates vary by model; see the pricing table.

The blended-rate trap

A single blended per-token price hides the input/output split. If your workload is heavy on output, a model with cheap input but expensive output will cost far more than the blended number suggests. Always price input and output separately against your real usage ratio rather than trusting one blended figure.

How to compare fairly

Fix the workload and vary only the model: same input tokens, same output tokens, same monthly volume. That isolates price from prompt differences. The compare tool does exactly this, and the pricing directory lists every rate with its official source and verification date.