0%
Applied AI Systems
Retrieval and Knowledge
Agents and Reasoning
Evaluation, Interpretability, and Safety
Production Research Engineering
Building a Code Agent: Case Study
A code agent is an LLM that can read a codebase, propose changes, run tests, and iterate until a bug is fixed or a feature is done. In 2026 this is the most demanding agent pattern in production. Devin, Claude Code, Cursor Agent, GitHub Copilot Workspace, Aider, OpenHands, and SWE-agent are all variants of the same basic loop, with different choices about tools, context management, and when to surrender back to the human.
This lesson is a case study. By the end you should be able to sketch a code agent on a whiteboard, name every component, and argue for its design. We will borrow shamelessly from the public architectures of Claude Code and SWE-agent, because the honest engineering trade-offs are only visible when you compare real systems.
What a Code Agent Actually Does
Start from the user story. A developer files a bug: "When I call reset_password with an empty email the server crashes instead of returning a 400." The developer does not want a chat explanation. They want a pull request that fixes the bug and includes a test. A code agent takes that issue and the repo and tries to produce that pull request.
The work breaks into five phases that every code agent performs, even if the surface looks like free-form chat:
- Understand the issue. Parse the bug report, identify the symptom and the suspected surface area.
- Navigate the repo. Find the files that implement
reset_passwordwithout reading the entire codebase. - Plan a fix. Decide what the change should look like in prose before writing any code.
- Edit and test. Apply a patch, run tests, observe failures, iterate.
- Summarize and hand off. Produce a commit message or PR description that a human can review in under a minute.
Every piece of that pipeline has at least one way to fail, and the agent's job is to detect failure and recover. Most production code agents spend more engineering effort on failure recovery than on the happy path.
The Minimum Tool Set
A code agent needs a very small number of tools. You can ship a working agent with four:
- read_file(path, start_line=None, end_line=None). Returns the text of a file, optionally a line range. This is the most-called tool and the source of most context-window pressure.
- write_file(path, content) or apply_patch(patch). Writes changes. We will argue in Section 3 that
apply_patchis the safer form. - run_shell(cmd, timeout_seconds). Runs a command in a sandbox. The agent uses this for tests, linters, type checkers, and ad-hoc scripts. One tool covers all of them because the model knows shell syntax.
- code_search(query, path_prefix=None). Grep or symbol search across the repo. Lets the agent find definitions and call sites without reading every file.
Claude Code ships a slightly larger set (edit, bash, glob, grep, read, todo), but every extra tool is a choice point in the model's head that can be made wrong. A tighter set is usually easier to prompt. SWE-agent's ACI (Agent-Computer Interface) paper is explicit about this: they found that giving the model too many ways to read a file made it pick the wrong one. The fix was to consolidate.
Tool count is a liability, not an asset. Every tool is a branch in the decision tree the model has to learn. The SWE-agent team reported that collapsing multiple file-read tools into one improved benchmark scores by several points. If you can express a capability by composing existing tools, do that before adding a new one.
The High-Level Loop
The agent loop is a pseudocode function with a fixed iteration budget. Everything else is details.
Look at this loop carefully. It contains every design decision in a code agent, bundled into function names you still need to implement:
read_relevant_files: the file navigation strategy (Section 2).llm.planandllm.write_patch: two separate calls, not one. Planning and coding are different skills and benefit from different prompts. This is the classic ReAct split applied to code.apply_patch: has to handle diff-application failures cleanly, because the model does generate unapplyable diffs (Section 3).run_tests: needs a sandbox, a timeout, and failure parsing.update_context: the hardest part, how to fold test failures back into the next planning prompt without blowing up the context window.max_iters: the giving-up condition. Without a budget, the agent will loop forever on flaky tests. Claude Code uses 25 as its default ceiling for a single sub-task. Higher budgets help on SWE-Bench scores but bankrupt the agent's cost model in production.
Actor, Critic, or Just Actor?
Some code agent papers use a two-model setup: an actor LLM that proposes edits and a critic LLM that reviews them before they run. OpenHands supports this pattern. In practice, teams have mostly dropped it. Why? Because running tests is a free, perfect critic. The compiler tells you if the code parses; the test suite tells you if it works; the linter tells you if it is clean. Nothing the LLM critic says is as reliable as "the tests pass."
The actor-critic pattern is more useful when feedback is slow, noisy, or absent, which is exactly the situation for research agents (Lesson 56) but not for code. This is one of the few places where code agents are easier than research agents: ground truth is cheap.
If you find yourself adding an LLM critic to a code agent, first check whether the tests, linter, or type checker can give you the same signal. Free critics beat learned critics whenever they exist, and code has more free critics than almost any other domain.
Real Systems for Context
Concrete anchors help. Here is how the major code agents in 2026 map onto this architecture:
- Claude Code (Anthropic). Terminal-based. Tools: read, edit, bash, glob, grep, todo. One actor, tight iteration loop, explicit todo list that the agent maintains as state. Aimed at developer workflows, not headless batch runs.
- Cursor Composer / Agent (Cursor). IDE-integrated. Gets a stream of file context from the editor. Very aggressive context gathering via code graph. Much shorter loop because the human is watching.
- Devin (Cognition). Long-horizon, autonomous. Runs in a cloud VM with a browser and a shell. Has a scratchpad and an explicit "knowledge" memory store across sessions. Designed for multi-hour tasks a human mostly ignores.
- SWE-agent (Princeton / academia). Optimized for SWE-Bench. Custom ACI with careful tool design. The paper is the best technical read in this space because the authors document their iteration.
- Aider. Terminal-based, older, conservative. Maps files to a git repo map, uses unified-diff edits, has a very explicit conversation structure. Its real contribution is the repo map idea, which we will see in Section 2.
- OpenHands (formerly OpenDevin). Open-source. Supports multiple agent backends. Good for experimenting with the loop itself.
All of these systems share the same core loop. They differ in three places: how they navigate files, how they format edits, and how they handle failure. Those are the three sections that follow.