Cost Modeling and Optimization

Topics Covered

Cost Components of an LLM System

Token Pricing in 2026

The Cost per Query Formula

Where the Money Actually Goes

Building a Cost Model from a Query Trace

The Unit Economics Question

Scaling Example: A 10K-User Chatbot

Caching Strategies

Exact-Match Caching

Semantic Caching

Prompt Caching (Prefix Caching)

When Caching Hurts

Layering the Three Cache Types

Model Routing and Cascades

The Cascade Pattern

The Cascade Decision Math

Confidence Threshold Routing in Practice

Router Systems in the Wild

Practical Tuning Tips

Prompt Compression

Where Compression Opportunities Live

LongLLMLingua-Style Compression

Summarization in the Loop

When Compression Hurts

Putting It Together

LLM products have a strange cost structure. A SaaS company built on a relational database can usually ignore per-query cost until they have a million users. An LLM company starts feeling cost pain at one hundred users. The reason is simple: every user interaction burns real tokens on someone else's GPU, and the bill arrives the following month with uncomfortable clarity.

Before you can optimize anything, you need to know where the money actually goes. Most teams guess, and most teams guess wrong. They blame the big model when the real culprit is a retrieval call that embeds the entire corpus on every query, or an evaluation harness that reruns the leaderboard twice a day, or a debug log that dumps full prompts into a vector store nobody remembers exists. This section walks through the full cost surface of an LLM application and shows you how to build a cost model from a query trace that tells you exactly where every cent goes.

This lesson pairs with the earlier inference_optimization lesson. That one covered the model-side cost: quantization, paged attention, speculative decoding, the levers the provider pulls to make a single GPU serve more requests. This one covers the application-side cost: caching, routing, prompt compression, the levers you pull as the engineer who built the product. You need both.

Token Pricing in 2026

Almost every hosted LLM is priced per token, with different rates for input and output. Input tokens are the prompt you send. Output tokens are the text the model generates. Output tokens are always more expensive than input tokens, often by a factor of four, because generation is sequential and forces the GPU to run one step at a time while input can be processed in a single batched pass.

Representative prices as of early 2026, per million tokens:

ModelInput $/MtokOutput $/MtokRatio
gpt-4o2.5010.004x
gpt-4o-mini0.150.604x
claude-3-5-sonnet3.0015.005x
claude-3-5-haiku0.804.005x
llama-3-70b (Together)0.880.881x
llama-3-8b (Groq)0.050.081.6x

Open-weight models hosted on Together, Groq, or Fireworks often have symmetric or near-symmetric pricing because the provider is selling GPU time directly, not packaging it into a tier with a target margin. This matters for routing decisions later: if you mostly generate long outputs, llama-3-70b on Together can undercut gpt-4o-mini on total cost even though gpt-4o-mini looks cheaper per input token.

The Cost per Query Formula

The headline number you care about is dollars per query. For a single model call, the math is straightforward:

costquery=nin⋅pin+nout⋅pout\text{cost}_{\text{query}} = n_{\text{in}} \cdot p_{\text{in}} + n_{\text{out}} \cdot p_{\text{out}}

where ninn_{\text{in}} is the input token count, noutn_{\text{out}} is the output token count, and pinp_{\text{in}}, poutp_{\text{out}} are the per-token prices. A 1000-input-token, 500-output-token query on gpt-4o costs 1000 times 2.5e-6 plus 500 times 10e-6, which is 0.0025 plus 0.005, or 0.0075 dollars. Three quarters of a cent. Looks tiny. Multiply by a million queries and you have 7,500 dollars per month for that single call pattern.

cost=nin⋅pin+nout⋅pout\text{cost} = n_{\text{in}} \cdot p_{\text{in}} + n_{\text{out}} \cdot p_{\text{out}}
python
1def query_cost(in_tokens, out_tokens, model='gpt-4o'):
2    prices = {
3        'gpt-4o':            (2.50e-6, 10.00e-6),
4        'gpt-4o-mini':       (0.15e-6,  0.60e-6),
5        'claude-3-5-sonnet': (3.00e-6, 15.00e-6),
6        'claude-3-5-haiku':  (0.80e-6,  4.00e-6),
7        'llama-3-70b':       (0.88e-6,  0.88e-6),
8    }
9    p_in, p_out = prices[model]
10    return in_tokens * p_in + out_tokens * p_out
11
12for m in ['gpt-4o', 'gpt-4o-mini', 'claude-3-5-sonnet', 'llama-3-70b']:
13    c = query_cost(1000, 500, m)
14    print(f'{m:20s} ${c:.5f} per query, ${c * 1_000_000:.0f} per million')
Token cost dominates at scale. A 100x cheaper model often costs 100x less per query if you can route the easy queries to it.

Where the Money Actually Goes

A production LLM system has more than one model call per user request. A typical RAG pipeline might look like this for a single user turn:

  1. Embed the user query with an embedding model.
  2. Run a vector similarity search in a managed vector database.
  3. Rerank the top candidates with a small cross-encoder.
  4. Call the main LLM with the retrieved context and the user query.
  5. Log the prompt, retrieved docs, and response to an evaluation store.
  6. Occasionally trigger an eval run that reprocesses a sample of traffic.

Each of those steps has a cost. Embedding models are cheap per token (roughly 0.02 to 0.10 dollars per million tokens) but you call them on every query and every document you index. Vector databases charge per read, per write, and per stored vector, which adds up fast when your corpus is in the millions of docs. Reranker calls are small LLM calls that you make many of per query. Eval runs are silent background expenses that teams routinely underestimate.

Priced on the same 1,500 in 400 out query, the six model table spans 98.1x rather than the over 20x claimed for the five models it names.

A realistic share-of-spend breakdown for a mid-sized RAG chatbot, measured on actual production traces from teams the author has worked with:

ComponentTypical share of total cost
Main LLM generation55-75%
Embedding calls (queries + indexing)5-15%
Vector DB reads and storage3-10%
Reranker LLM calls5-15%
Logging and evaluation pipelines5-20%

Evaluation can surprise you. A team running nightly evals on 10,000 examples against three candidate models can easily spend more on evals than on user-facing inference during the first few months of iteration. This is not wrong, it is investment in knowing your system, but it needs to be a conscious decision, not an invisible line item.

Building a Cost Model from a Query Trace

The only reliable way to understand your cost structure is to instrument it. For every user request, record the full chain of model calls with their input and output token counts, plus the model used. Store that in a columnar store, or just a SQLite file if you are small, and run SQL queries against it.

A minimal cost model has three tables or derived columns:

  • calls: one row per LLM call, with trace_id, component (embed, retrieve, rerank, generate), model, input_tokens, output_tokens, timestamp.
  • prices: current price per million input and output tokens per model, versioned by date so old traces use their historical prices.
  • traces: one row per user-facing request, with total cost computed as the sum of calls joined to prices.

Once you have this, the questions you usually could not answer become trivial:

  • What is the cost of a median user turn? A p99 user turn? (Often 10x apart.)
  • Which component dominates cost for each persona? (Chat-heavy vs search-heavy users have very different cost profiles.)
  • What percentage of total spend goes to the longest 1% of conversations?
  • How does cost trend week over week, and is it tracking signup growth or something else?
Key Insight

Before optimizing anything, trace one week of production traffic and rank components by total cost. The biggest cost line is almost never the one the team expects. We once found a RAG system spending 40 percent of its budget on a debug endpoint that embedded the entire prompt twice for a deprecated dashboard. Nobody had looked in months.

The Unit Economics Question

Every LLM product eventually faces the same question: does the cost per query leave room for a profitable business? If you charge 20 dollars per month and the average active user makes 1,000 queries at 0.01 dollars each, you are spending 10 dollars of LLM cost per 20 dollar subscription. That leaves 10 dollars for everything else: engineers, infrastructure, sales, support, profit. It is not hopeless but it is uncomfortable.

The same product with cost per query cut to 0.001 dollars (10x reduction) leaves 19 dollars of margin. That is the difference between a business and a treadmill. Every technique in the rest of this lesson is about moving from the first scenario to the second, ideally without the user noticing any quality change.

Eugene Yan's blog post "Patterns for Building LLM-based Systems and Products" is the canonical field guide here. He catalogs the techniques used by teams that have been running LLM products at scale for years. If you read only one external reference on this topic, read his.

Scaling Example: A 10K-User Chatbot

Let us walk through concrete numbers for a realistic scale. Imagine a product support chatbot with 10,000 monthly active users, each averaging 100 queries per month. That is 1,000,000 queries per month. Suppose each query involves:

  • 1,500 input tokens (user query plus retrieved context)
  • 400 output tokens (response)
  • One embedding call for the user query (50 tokens)
  • One vector DB read (negligible for this exercise)

For the main generation call on different models, using 1500 in and 400 out:

  • gpt-4o: 1500 times 2.5e-6 plus 400 times 10e-6 equals 0.00775 per query, or 7,750 per month.
  • gpt-4o-mini: 1500 times 0.15e-6 plus 400 times 0.6e-6 equals 0.000465 per query, or 465 per month.
  • claude-3-5-sonnet: 1500 times 3e-6 plus 400 times 15e-6 equals 0.0105 per query, or 10,500 per month.
  • claude-3-5-haiku: 1500 times 0.8e-6 plus 400 times 4e-6 equals 0.0028 per query, or 2,800 per month.
  • llama-3-70b (Together): 1500 times 0.88e-6 plus 400 times 0.88e-6 equals 0.001672 per query, or 1,672 per month.

The range is over 20x between the cheapest and most expensive choice for the same workload. If your average user pays 10 dollars a month, a 7,750 dollar generation bill on 10,000 users gives you 92,250 left for everything else, which is fine. At 100,000 users on the same model, generation alone is 77,500 dollars, and at that scale evaluation, embedding, and vector DB costs also multiply. The numbers start to bite.

The optimization question is not "which single model is cheapest." It is "how do I get 95 percent of the quality of the expensive model at 10 percent of the cost." The rest of this lesson answers that question with caching, routing, and compression.