0%
Applied AI Systems
Retrieval and Knowledge
Agents and Reasoning
Production Research Engineering
Production Operations and Safety
Multi-Agent Systems and Case Studies
Mechanistic Interpretability
Mechanistic interpretability is the project of understanding what is happening inside neural networks. Not just what they do, but how they do it. It treats neural networks as objects to be reverse-engineered, with the goal of discovering the algorithms they have learned. This lesson covers the foundational techniques: probing, induction heads, superposition, and activation patching.
Why Interpretability Matters
Neural networks are notoriously opaque. We can train them, evaluate them, and observe their outputs, but the internal computations that produce those outputs are mostly mysterious. This is uncomfortable for several reasons:
- Safety: We cannot trust systems we do not understand. As LLMs are deployed in more critical applications, understanding how they make decisions becomes a safety prerequisite.
- Debugging: When a model fails, knowing why is hard without understanding the internals. Interpretability tools could turn debugging from guesswork into science.
- Scientific curiosity: Neural networks are doing something computationally interesting. Understanding that might teach us about intelligence in general.
- Regulation: As AI systems are subject to legal and regulatory frameworks, the ability to explain decisions becomes important for compliance.
Mechanistic interpretability is the most ambitious approach to these problems. Instead of treating the model as a black box and making post-hoc explanations, it tries to understand the actual circuits inside the network.
Beginner: understand what probing is and why correlation is not causation.
Intermediate: be able to design a probing experiment with proper controls and explain superposition.
Advanced: know when to use activation patching vs. sparse autoencoders and reason about circuit-level analysis.
Probing Classifiers: The Starting Point
Probing is the simplest interpretability technique and a good entry point. The idea: take a model's internal activations at some layer, train a small classifier to predict some property from those activations, and use the classifier's accuracy as evidence of what the model "knows" at that layer.
For example, suppose you want to know whether a language model represents grammatical features like part-of-speech. You take the model's hidden states at a specific layer, train a small linear classifier to predict part-of-speech tags from those hidden states, and measure accuracy on held-out data. High accuracy suggests the part-of-speech information is present at that layer.
Probes are widely used to study what kinds of information are present in neural network representations. Researchers have probed for syntactic features, semantic categories, factual knowledge, world knowledge, mathematical concepts, and many other properties. The general finding is that LLMs encode a remarkable amount of information in their hidden states, often in ways that are not obvious from their input-output behavior.
The Probing Caveat
Probing has a fundamental limitation: probe accuracy does not prove the model uses that information. A probe could pick up information that is encoded in the activations but never used by downstream computations. The model might "know" the part-of-speech tag in some abstract sense without that knowledge being part of how it produces outputs.
This is the "selectivity vs. accuracy" problem. A more accurate probe might just be detecting incidental correlations rather than functional structure. To address this, researchers use techniques like control tasks (probe a randomized version of the data), smaller probes (limit the probe's capacity to learn correlations), and causal interventions (modify the activations and see if behavior changes).
Causal probing, which we will discuss with activation patching later, is more rigorous than correlational probing because it tests whether the information is actually used.
What Probes Have Found
Despite the limitations, probing has produced many useful insights about what LLMs encode:
- Linguistic features: Models learn syntactic categories, dependency structures, and morphological features without being explicitly trained on them.
- Factual knowledge: Specific facts can be located in specific layers and even specific neurons.
- World models: Some models encode spatial, temporal, and physical relationships that go beyond surface text patterns.
- Truth and falsity: Models often have internal representations that distinguish true from false statements, even when they generate false statements.
- Calibration signals: Models often "know" when they are uncertain, even when their outputs do not express the uncertainty.
The last finding is particularly interesting for safety. If models internally represent uncertainty better than they express it, there may be ways to extract calibrated confidence directly from internal representations, bypassing the model's potentially miscalibrated verbal output.
Models often 'know' more than they say. Probing reveals internal representations of truth, uncertainty, and factual knowledge that do not always surface in outputs. This gap between internal knowledge and external behavior is one of the most important findings in interpretability research.
Linear Probes vs. Nonlinear Probes
Probes can be linear or nonlinear. Linear probes are simple, they fit a single matrix from activations to predictions. Nonlinear probes (small MLPs) are more expressive but can also learn more incidental correlations.
The convention in interpretability research is to prefer linear probes when possible. The reasoning is that if a feature is present in a linearly extractable form, it is more likely to be the kind of feature the model itself can use, since the model's downstream layers operate in linear-then-nonlinear patterns. Nonlinear features could be artifacts of the probe's expressiveness, not the model's representation.
The "linear representation hypothesis" states that meaningful features in LLMs are encoded as directions in activation space, accessible by linear projections. This hypothesis is partially supported by empirical evidence and is the foundation for many interpretability techniques.
Probing in Practice
Practical probing has become a standard workflow. Researchers extract activations from a target layer, train a probe, and report results. Tools like TransformerLens make this easy for popular LLM architectures.
The findings inform other interpretability work. If a probe finds a specific feature at a specific layer, that becomes a hypothesis for circuit-level analysis: "the model represents X at layer N, so the next step is to find which attention heads or MLP neurons compute X". Probing identifies what to look for; circuit analysis finds the mechanisms.