0%
Applied AI Systems
Retrieval and Knowledge
Agents and Reasoning
Production Research Engineering
Production Operations and Safety
Multi-Agent Systems and Case Studies
LLM Evaluation
Evaluation is the part of LLM development that nobody is satisfied with. Every benchmark gets gamed eventually. Every metric eventually fails to track what we actually care about. And yet, evaluation is the only way to know whether progress is real. This lesson covers the major approaches to LLM evaluation, why each is broken, and how to think about evaluation in practice.
Why Benchmarks Exist
Before benchmarks, LLM progress was an anecdote economy. Researchers would say "this model is better" and you would have to take their word for it. Benchmarks gave the field a shared currency: a set of fixed test inputs with known correct answers that any model could be evaluated against. MMLU, HellaSwag, ARC, GSM8K, HumanEval, BIG-Bench. These names define modern LLM evaluation.
A good benchmark has three properties. It is representative, performance correlates with real-world utility. It is stable, the same model gets similar scores across runs. And it is uncontaminated, no model has seen the test data during training. The first two are hard. The third is the central scandal of LLM evaluation.
What Contamination Looks Like
Contamination happens when test data leaks into training data. The model learns the test answers directly rather than the underlying capability. A contaminated benchmark gives high scores that do not predict real-world performance.
Contamination has many sources. Benchmarks are often hosted on GitHub or arXiv, and pretraining corpora include GitHub and arXiv. When researchers release a new benchmark, it is on the open web within days, and the next pretraining run sees it. Even paraphrased or translated versions of the benchmark contaminate, because the model can pattern-match the structure.
The problem compounds with model size. Larger models have more capacity to memorize, and large pretraining corpora are more likely to contain test data. By the time GPT-4 was released, most popular benchmarks were partially or fully contaminated.
Detecting Contamination
Several techniques exist to detect contamination, but none are perfect:
- String matching: Search the training corpus for exact matches of test questions. Catches direct copies but misses paraphrases.
- Perplexity gap: Compare the model's perplexity on test questions versus paraphrased or shuffled versions. A large gap suggests memorization.
- Membership inference attacks: Use techniques from privacy research to infer whether specific examples were in training data.
- Held-out splits: Hold a portion of the benchmark private. If the public and private scores diverge, the public split is contaminated.
- Date-based comparison: Test the model on examples from after its training cutoff. If old examples perform much better than new ones, contamination is likely.
None of these provides a clean signal. Researchers report contamination rates ranging from 1% to 50% on the same benchmark depending on the detection method. The honest answer is that no one knows exactly how contaminated current benchmarks are.
Living Benchmarks
The response to contamination is "living benchmarks", evaluation sets that change over time so memorization cannot help. Examples include:
- LMSYS Arena: Uses live human preference judgments on novel prompts. The "test set" is whatever users type. Contamination is impossible because there is no fixed test set.
- GPQA Diamond: A small set of expert-written questions in physics, chemistry, and biology. Designed to be hard enough that it remains useful after contamination because solving requires actual reasoning.
- SWE-Bench: Real GitHub issues that models must fix. New issues are added over time, so the benchmark evolves.
- AIME: Math olympiad problems that are released each year. Models can be evaluated on the most recent year's problems, which are guaranteed to be after training cutoffs.
Living benchmarks have their own problems, they are harder to standardize, results are noisier, and they require ongoing curation. But they are the current best answer to contamination.
Contamination is not always intentional. Benchmarks get published on GitHub and arXiv, which end up in pretraining corpora. The model memorizes answers without anyone planning it. This is why contamination detection is part of responsible evaluation, not just an academic concern.
What Benchmarks Cannot Measure
Even uncontaminated benchmarks miss important capabilities. They typically measure:
- Knowledge recall (multiple choice over facts)
- Reasoning over short, well-defined problems
- Code generation on small tasks
- Instruction following on isolated prompts
They typically miss:
- Long-context understanding
- Multi-turn conversation quality
- Tool use and agentic behavior
- Calibration and uncertainty
- Robustness to distribution shift
- Safety and refusal behavior
- Real-world utility on novel tasks
The result is that benchmark scores can dramatically outpace real-world performance. A model that scores 90% on MMLU might be barely better than a 75% model on tasks users actually care about. This is one reason the field has grown skeptical of "SOTA on benchmark X" as a meaningful claim.