LLM A/B Testing and Experimentation

Topics Covered

Experiment Design for LLMs

What Is the Variant?

Randomization Unit

The Carryover Problem

Power Analysis Before You Ship

Picking the Primary Metric

Online Evaluation With Real Traffic

Implicit Signals: What Users Do

Explicit Signals: What Users Say

Tooling: What Actually Logs This Stuff

Instrumenting the Right Events

Avoiding Survivorship Bias

Offline Eval vs Online Eval

When Offline Matches Online

When Offline Diverges From Online

The 5 Percent Offline, 0 Percent Online Pattern

The Recall vs Quality Gap

Using Offline Eval to Filter Candidates for Online

Building an Offline Set That Tracks Online

The Cost Tradeoff

Statistical Significance for Subjective Metrics

Why t-tests Are Often Wrong for LLM Metrics

Mann-Whitney U for Ordinal Feedback

Bootstrap Confidence Intervals for Thumbs Ratios

Effect Size Matters as Much as p-values

The Multiple-Comparison Trap

The Peeking Problem

Putting It Together

Running an A/B test on a conventional product feature is a mature discipline. You flip a flag on for half your users, measure a metric, compare the two groups. Frameworks like Statsig, Eppo, and GrowthBook package all the machinery: variant assignment, exposure logging, SRM checks, stratified analysis. LLM systems break enough of the assumptions behind those frameworks that you have to rethink experiment design from first principles. Not because the statistics are different, but because what counts as a "variant" and what counts as a "unit" are much fuzzier when the product is a model that generates free text.

This section is about the setup decisions. If you get them wrong, no amount of clever analysis at the end will rescue the result.

What Is the Variant?

In a traditional product experiment, the variant is a deliberate code change: new button color, different pricing, alternate onboarding flow. For an LLM system, the space of possible variants is much richer, and the decision of what to vary is itself the first design question. Think of variants as living at four levels of a stack, and the level determines almost everything downstream (cost, sensitivity, interpretation).

Model variants swap one underlying LLM for another. gpt-4o-mini versus gpt-4o, or claude-3.5-sonnet versus claude-3.5-haiku, or an open-source Llama-3.1-70B versus a hosted Anthropic model. These are the highest-leverage changes: a different model can change latency by 5x, cost by 10x, and quality across dozens of subtle dimensions simultaneously. They are also the hardest to diagnose, because a regression could come from any one of those dimensions.

Prompt variants keep the model fixed and change the system prompt, few-shot examples, or instruction phrasing. This is where most day-to-day experimentation happens because prompts are cheap to iterate. A product team might run 5-10 prompt A/B tests per week during active development. The downside is that prompt changes often have tiny effect sizes — a 2-3 percent lift on implicit signals — and tiny effects require large samples to detect.

Retrieval variants change what context the model sees: a different embedding model, a new chunking strategy, a re-ranker added to the pipeline, a bumped top_k. These variants can produce dramatic quality improvements on grounded tasks, but they also interact with the prompt: changing retrieval without updating the prompt to match the new context shape can leave quality flat.

Pipeline variants change the overall architecture: single-shot versus multi-turn, adding a critic step, using a router between two models. These are rare but high-stakes. They tend to ship as entire product releases rather than as A/B tests, because they change latency, cost, and failure modes in ways that make a clean comparison difficult.

Key Insight

A rule I follow: never vary two levels of the stack in one experiment. If you change both the model and the prompt, you have confounded the two sources of lift and will not know which one to credit. Run the model swap with the old prompt first, then separately run a prompt experiment with the new model. It is slower but it actually lets you learn.

Randomization Unit

The second design decision is the randomization unit — the thing you flip a coin for to decide which variant a user sees. In a web A/B test the unit is almost always the user. In an LLM system you have three plausible choices and each has a different bias-variance tradeoff.

User-level randomization assigns a variant to a user ID (or device fingerprint for anonymous traffic) and sticks with it across sessions. This is the cleanest design because a single user always sees a single variant, which removes within-user contamination. The downside is that the variance between users is high: some users ask ten questions a day, others ask ten a year, and the heavy-tailed distribution of user activity means a handful of power users can dominate your metric. You need more users to hit a given power target than you would with a more granular unit.

Session-level randomization re-randomizes at the start of each session. A session can be defined as a contiguous block of activity (e.g., no gap longer than 30 minutes). Sessions are more numerous than users, which improves sensitivity, and users do not remember variant assignments across sessions for long enough to cause confusion. The risk is that the same user can bounce between variants and form expectations — "the morning assistant was better than the afternoon one" — that leak across the boundary.

Query-level randomization flips a coin for every single request. This is the most sensitive design (many more independent units) but the carryover problem becomes severe: the user sees variant A answer question 1, then variant B answer question 2, and those answers may look inconsistent or contradict each other. Users notice. They complain. They write angry tweets about the AI being "schizophrenic." For any product with stateful conversations or any task where consistency matters, query-level randomization is usually ruled out.

A two point lift needs 3,600 sessions an arm, which is 24 days at 300 a day, and the randomization unit you pick moves that calendar by an order of magnitude.

For most LLM products, session-level randomization hits the sweet spot: enough independence to power small effects, few enough switch points to avoid visible inconsistency. User-level is the fallback for products where a single session blends multiple tasks and switching is disorienting.

The Carryover Problem

Carryover is the term for treatment effects that leak across randomization units. In a traditional web experiment, carryover is rare: if I see variant B of the homepage today and variant A tomorrow, I probably do not remember enough to be biased by it. LLM systems are different in two ways that make carryover a real threat.

First, users develop mental models of the assistant quickly. After a dozen interactions they form opinions about "how it talks" and "what it is good at," and those opinions carry across sessions. A user who had a great day with variant B may rate variant A harshly the next day because the comparison is visible to them, even though the experimental design says it should not be.

Second, LLM outputs often feed back into the user's context. A user who asks a question, gets a good answer, and copies it into a document has been trained by the assistant. If they then ask a follow-up question with different phrasing (because the first answer taught them the vocabulary), their behavior on the new variant is no longer an unbiased sample from the "natural query distribution." This is especially bad for query-level randomization where the user's own queries become test stimuli that were shaped by the other arm.

Mitigations: prefer session-level or user-level randomization to minimize switch points; run experiments long enough that the wash-in period (when users are forming expectations) gets swamped by the steady-state period; and report results on the post-wash-in window, not on day 1. A common heuristic is to discard the first 2-3 days of data and analyze only the stable portion.

Power Analysis Before You Ship

Most teams skip power analysis and then discover, two weeks into an experiment, that the effect they care about is too small to detect with the sample they have. The LLM version of this mistake is especially painful because LLM experiments tend to have small effect sizes on noisy ordinal metrics. You should always run a power calculation before starting. For a two-sample comparison of a continuous metric, the required sample size per arm is roughly n=16σ2/δ2n = 16 \sigma^2 / \delta^2 at 80 percent power and a two-sided significance level of 0.05, where σ\sigma is the metric standard deviation and δ\delta is the minimum detectable effect.

Concretely: if your implicit-feedback score has a standard deviation of 0.3 on a 0-1 scale and you want to detect a 0.02 (2 percentage point) lift, you need about n=16⋅0.09/0.0004=3600n = 16 \cdot 0.09 / 0.0004 = 3600 samples per arm. That sounds small until you remember that a "sample" here is probably a session, not a query, and that sessions are rarer than queries. For products with tens of thousands of sessions per day, 3600 per arm is an overnight experiment. For products with hundreds of sessions per day, it is two weeks minimum. Many LLM features sit in the second regime during early development, which is why experiments on small products often look underpowered.

Picking the Primary Metric

The final design decision is the primary metric — the one metric you commit to before the experiment begins and use to make the ship/no-ship call. Picking it well requires resisting two temptations: picking a metric that is easy to move (but does not reflect user value) and picking a "god metric" that captures everything (but moves slowly and noisily). A good primary metric is sensitive to real quality changes, resistant to gaming, and cheap enough to compute that you get daily updates. Examples that work in practice: thumbs-up rate on the first response of a session, session completion rate, percent of sessions with at least one copy-paste event, percent of sessions without a regenerate click. Examples that do not work: NPS (too slow), subscription conversion (too many confounders), LLM-judge score (too expensive to compute at scale and not truly user-grounded).

Eugene Yan's blog series "Evals for LLMs" makes this point repeatedly: the best primary metrics for online LLM experiments are implicit behavioral signals, not explicit ratings. Pick one, commit to it, and do not change it mid-experiment.