What prompt caching is
Prompt caching lets a provider store a prompt prefix it has seen recently and bill it at a discounted cached-input rate on later calls. Instead of paying the full input price for a long system prompt on every request, you pay full price once and the discounted rate after that, for as long as the prefix stays unchanged.
When it pays off
Caching helps when two conditions hold: the prefix is large, and it repeats across many requests. A long, fixed system prompt or a stable block of retrieved context that every request shares is the classic win. A short prompt, or one that changes on every call, gains little because there is little to reuse.
How much you can save
Savings depend on what fraction of your input is cacheable and how big the cached discount is. Estimate the cacheable share (usually the system prompt) and the prompt caching calculator applies the cached rate to that fraction for you. Background on the token types is in the input, output and cached tokens guide.