The anatomy of a RAG bill

Retrieval-augmented generation adds a retrieval step before the model answers: a query fetches relevant chunks, those chunks are prepended to the prompt, and the model reads all of it. The bill is driven by tokens per query (system prompt, retrieved chunks, the question, and the answer) multiplied by query volume.

Why context dominates

The retrieved context is usually the largest input component, and it is billed on every single query. A question that retrieves 3,000 tokens of chunks sends those 3,000 tokens each time, even if the question itself is short. This is why input cost, not output cost, tends to dominate a RAG workload.

How to cut RAG cost

Retrieve fewer, better chunks. Tightening retrieval so only the most relevant context is sent cuts input tokens directly and often improves answer quality at the same time. Summarizing or compressing chunks, and caching any stable prefix, help further. Model the numbers on the RAG cost calculator.