0%
Applied AI Systems
Retrieval and Knowledge
Agents and Reasoning
Evaluation, Interpretability, and Safety
Production Research Engineering
Multi-Agent Systems and Case Studies
Cost Modeling and Optimization
LLM products have a strange cost structure. A SaaS company built on a relational database can usually ignore per-query cost until they have a million users. An LLM company starts feeling cost pain at one hundred users. The reason is simple: every user interaction burns real tokens on someone else's GPU, and the bill arrives the following month with uncomfortable clarity.
Before you can optimize anything, you need to know where the money actually goes. Most teams guess, and most teams guess wrong. They blame the big model when the real culprit is a retrieval call that embeds the entire corpus on every query, or an evaluation harness that reruns the leaderboard twice a day, or a debug log that dumps full prompts into a vector store nobody remembers exists. This section walks through the full cost surface of an LLM application and shows you how to build a cost model from a query trace that tells you exactly where every cent goes.
This lesson pairs with the earlier inference_optimization lesson. That one covered the model-side cost: quantization, paged attention, speculative decoding, the levers the provider pulls to make a single GPU serve more requests. This one covers the application-side cost: caching, routing, prompt compression, the levers you pull as the engineer who built the product. You need both.
Token Pricing in 2026
Almost every hosted LLM is priced per token, with different rates for input and output. Input tokens are the prompt you send. Output tokens are the text the model generates. Output tokens are always more expensive than input tokens, often by a factor of four, because generation is sequential and forces the GPU to run one step at a time while input can be processed in a single batched pass.
Representative prices as of early 2026, per million tokens:
| Model | Input $/Mtok | Output $/Mtok | Ratio |
| gpt-4o | 2.50 | 10.00 | 4x |
| gpt-4o-mini | 0.15 | 0.60 | 4x |
| claude-3-5-sonnet | 3.00 | 15.00 | 5x |
| claude-3-5-haiku | 0.80 | 4.00 | 5x |
| llama-3-70b (Together) | 0.88 | 0.88 | 1x |
| llama-3-8b (Groq) | 0.05 | 0.08 | 1.6x |
Open-weight models hosted on Together, Groq, or Fireworks often have symmetric or near-symmetric pricing because the provider is selling GPU time directly, not packaging it into a tier with a target margin. This matters for routing decisions later: if you mostly generate long outputs, llama-3-70b on Together can undercut gpt-4o-mini on total cost even though gpt-4o-mini looks cheaper per input token.
The Cost per Query Formula
The headline number you care about is dollars per query. For a single model call, the math is straightforward:
where is the input token count, is the output token count, and , are the per-token prices. A 1000-input-token, 500-output-token query on gpt-4o costs 1000 times 2.5e-6 plus 500 times 10e-6, which is 0.0025 plus 0.005, or 0.0075 dollars. Three quarters of a cent. Looks tiny. Multiply by a million queries and you have 7,500 dollars per month for that single call pattern.
Where the Money Actually Goes
A production LLM system has more than one model call per user request. A typical RAG pipeline might look like this for a single user turn:
- Embed the user query with an embedding model.
- Run a vector similarity search in a managed vector database.
- Rerank the top candidates with a small cross-encoder.
- Call the main LLM with the retrieved context and the user query.
- Log the prompt, retrieved docs, and response to an evaluation store.
- Occasionally trigger an eval run that reprocesses a sample of traffic.
Each of those steps has a cost. Embedding models are cheap per token (roughly 0.02 to 0.10 dollars per million tokens) but you call them on every query and every document you index. Vector databases charge per read, per write, and per stored vector, which adds up fast when your corpus is in the millions of docs. Reranker calls are small LLM calls that you make many of per query. Eval runs are silent background expenses that teams routinely underestimate.
A realistic share-of-spend breakdown for a mid-sized RAG chatbot, measured on actual production traces from teams the author has worked with:
| Component | Typical share of total cost |
| Main LLM generation | 55-75% |
| Embedding calls (queries + indexing) | 5-15% |
| Vector DB reads and storage | 3-10% |
| Reranker LLM calls | 5-15% |
| Logging and evaluation pipelines | 5-20% |
Evaluation can surprise you. A team running nightly evals on 10,000 examples against three candidate models can easily spend more on evals than on user-facing inference during the first few months of iteration. This is not wrong, it is investment in knowing your system, but it needs to be a conscious decision, not an invisible line item.
Building a Cost Model from a Query Trace
The only reliable way to understand your cost structure is to instrument it. For every user request, record the full chain of model calls with their input and output token counts, plus the model used. Store that in a columnar store, or just a SQLite file if you are small, and run SQL queries against it.
A minimal cost model has three tables or derived columns:
calls: one row per LLM call, withtrace_id,component(embed, retrieve, rerank, generate),model,input_tokens,output_tokens,timestamp.prices: current price per million input and output tokens per model, versioned by date so old traces use their historical prices.traces: one row per user-facing request, with total cost computed as the sum ofcallsjoined toprices.
Once you have this, the questions you usually could not answer become trivial:
- What is the cost of a median user turn? A p99 user turn? (Often 10x apart.)
- Which component dominates cost for each persona? (Chat-heavy vs search-heavy users have very different cost profiles.)
- What percentage of total spend goes to the longest 1% of conversations?
- How does cost trend week over week, and is it tracking signup growth or something else?
Before optimizing anything, trace one week of production traffic and rank components by total cost. The biggest cost line is almost never the one the team expects. We once found a RAG system spending 40 percent of its budget on a debug endpoint that embedded the entire prompt twice for a deprecated dashboard. Nobody had looked in months.
The Unit Economics Question
Every LLM product eventually faces the same question: does the cost per query leave room for a profitable business? If you charge 20 dollars per month and the average active user makes 1,000 queries at 0.01 dollars each, you are spending 10 dollars of LLM cost per 20 dollar subscription. That leaves 10 dollars for everything else: engineers, infrastructure, sales, support, profit. It is not hopeless but it is uncomfortable.
The same product with cost per query cut to 0.001 dollars (10x reduction) leaves 19 dollars of margin. That is the difference between a business and a treadmill. Every technique in the rest of this lesson is about moving from the first scenario to the second, ideally without the user noticing any quality change.
Eugene Yan's blog post "Patterns for Building LLM-based Systems and Products" is the canonical field guide here. He catalogs the techniques used by teams that have been running LLM products at scale for years. If you read only one external reference on this topic, read his.
Scaling Example: A 10K-User Chatbot
Let us walk through concrete numbers for a realistic scale. Imagine a product support chatbot with 10,000 monthly active users, each averaging 100 queries per month. That is 1,000,000 queries per month. Suppose each query involves:
- 1,500 input tokens (user query plus retrieved context)
- 400 output tokens (response)
- One embedding call for the user query (50 tokens)
- One vector DB read (negligible for this exercise)
For the main generation call on different models, using 1500 in and 400 out:
- gpt-4o: 1500 times 2.5e-6 plus 400 times 10e-6 equals 0.00775 per query, or 7,750 per month.
- gpt-4o-mini: 1500 times 0.15e-6 plus 400 times 0.6e-6 equals 0.000465 per query, or 465 per month.
- claude-3-5-sonnet: 1500 times 3e-6 plus 400 times 15e-6 equals 0.0105 per query, or 10,500 per month.
- claude-3-5-haiku: 1500 times 0.8e-6 plus 400 times 4e-6 equals 0.0028 per query, or 2,800 per month.
- llama-3-70b (Together): 1500 times 0.88e-6 plus 400 times 0.88e-6 equals 0.001672 per query, or 1,672 per month.
The range is over 20x between the cheapest and most expensive choice for the same workload. If your average user pays 10 dollars a month, a 7,750 dollar generation bill on 10,000 users gives you 92,250 left for everything else, which is fine. At 100,000 users on the same model, generation alone is 77,500 dollars, and at that scale evaluation, embedding, and vector DB costs also multiply. The numbers start to bite.
The optimization question is not "which single model is cheapest." It is "how do I get 95 percent of the quality of the expensive model at 10 percent of the cost." The rest of this lesson answers that question with caching, routing, and compression.