Per-token pricing
LLM providers bill for usage, not seats. You pay for the tokens you send and the tokens the model generates. Prices are quoted per million tokens (sometimes per thousand). Because every request consumes tokens, your bill scales with total tokens processed, not with how many users you have.
Input, output and cached rates
There are three separately-priced rates. Input is your prompt and any resent context. Output is what the model generates, priced higher because generation is the expensive part. Cached input is a reused prefix billed at a discount where the provider supports it. The input, output and cached tokens guide covers each in detail.
The blended-rate trap
A single blended per-token price hides the input/output split. If your workload is heavy on output, a model with cheap input but expensive output will cost far more than the blended number suggests. Always price input and output separately against your real usage ratio rather than trusting one blended figure.
How to compare fairly
Fix the workload and vary only the model: same input tokens, same output tokens, same monthly volume. That isolates price from prompt differences. The compare tool does exactly this, and the pricing directory lists every rate with its official source and verification date.