Building a Code Agent: Case Study

Topics Covered

Requirements and Architecture

What a Code Agent Actually Does

The Minimum Tool Set

The High-Level Loop

Actor, Critic, or Just Actor?

Real Systems for Context

Planning and File Navigation

The File Map Approach

The RAG Approach

Building Context Within Token Limits

The Planning Prompt

When to Give Up and Ask

The Edit-and-Test Loop

Diff Format vs Full File Rewrite

The Patch Prompt

Applying Edits Safely

Running Tests

Turning Failures into Actionable Feedback

When to Revert

Known Failure Modes, Named

Evaluation on SWE-Bench

What SWE-Bench Is

SWE-Bench vs SWE-Bench Verified

Setting Up the Benchmark

What Scores Actually Mean

Metrics Beyond Pass Rate

The Gap Between Benchmark and Product

What a 20% Pass Rate Looks Like in Production

External References

A code agent is an LLM that can read a codebase, propose changes, run tests, and iterate until a bug is fixed or a feature is done. In 2026 this is the most demanding agent pattern in production. Devin, Claude Code, Cursor Agent, GitHub Copilot Workspace, Aider, OpenHands, and SWE-agent are all variants of the same basic loop, with different choices about tools, context management, and when to surrender back to the human.

This lesson is a case study. By the end you should be able to sketch a code agent on a whiteboard, name every component, and argue for its design. We will borrow shamelessly from the public architectures of Claude Code and SWE-agent, because the honest engineering trade-offs are only visible when you compare real systems.

What a Code Agent Actually Does

Start from the user story. A developer files a bug: "When I call reset_password with an empty email the server crashes instead of returning a 400." The developer does not want a chat explanation. They want a pull request that fixes the bug and includes a test. A code agent takes that issue and the repo and tries to produce that pull request.

The work breaks into five phases that every code agent performs, even if the surface looks like free-form chat:

  1. Understand the issue. Parse the bug report, identify the symptom and the suspected surface area.
  2. Navigate the repo. Find the files that implement reset_password without reading the entire codebase.
  3. Plan a fix. Decide what the change should look like in prose before writing any code.
  4. Edit and test. Apply a patch, run tests, observe failures, iterate.
  5. Summarize and hand off. Produce a commit message or PR description that a human can review in under a minute.

Every piece of that pipeline has at least one way to fail, and the agent's job is to detect failure and recover. Most production code agents spend more engineering effort on failure recovery than on the happy path.

One iteration is two model calls carrying the whole working set, so the 25th costs 2.80x the first and a budget 25 run spends 4.56 dollars before a single solve.

The Minimum Tool Set

A code agent needs a very small number of tools. You can ship a working agent with four:

  • read_file(path, start_line=None, end_line=None). Returns the text of a file, optionally a line range. This is the most-called tool and the source of most context-window pressure.
  • write_file(path, content) or apply_patch(patch). Writes changes. We will argue in Section 3 that apply_patch is the safer form.
  • run_shell(cmd, timeout_seconds). Runs a command in a sandbox. The agent uses this for tests, linters, type checkers, and ad-hoc scripts. One tool covers all of them because the model knows shell syntax.
  • code_search(query, path_prefix=None). Grep or symbol search across the repo. Lets the agent find definitions and call sites without reading every file.

Claude Code ships a slightly larger set (edit, bash, glob, grep, read, todo), but every extra tool is a choice point in the model's head that can be made wrong. A tighter set is usually easier to prompt. SWE-agent's ACI (Agent-Computer Interface) paper is explicit about this: they found that giving the model too many ways to read a file made it pick the wrong one. The fix was to consolidate.

Key Insight

Tool count is a liability, not an asset. Every tool is a branch in the decision tree the model has to learn. The SWE-agent team reported that collapsing multiple file-read tools into one improved benchmark scores by several points. If you can express a capability by composing existing tools, do that before adding a new one.

The High-Level Loop

The agent loop is a pseudocode function with a fixed iteration budget. Everything else is details.

patcht+1=actor(repo,testst,patcht,tracet)\text{patch}_{t+1} = \text{actor}(\text{repo}, \text{tests}_t, \text{patch}_t, \text{trace}_t)
python
1def code_agent(issue, repo_path, max_iters=10):
2    context = read_relevant_files(issue, repo_path)
3    for i in range(max_iters):
4        plan = llm.plan(issue, context)
5        patch = llm.write_patch(plan, context)
6        apply_patch(patch, repo_path)
7        result = run_tests(repo_path)
8        if result.passed:
9            return patch
10        context = update_context(context, patch, result.failure_trace)
11    return None  # gave up
A code agent's loop is read code, plan, patch, test, repeat. The patch+test loop is the inner cycle; file navigation is the outer cycle.

Look at this loop carefully. It contains every design decision in a code agent, bundled into function names you still need to implement:

  • read_relevant_files: the file navigation strategy (Section 2).
  • llm.plan and llm.write_patch: two separate calls, not one. Planning and coding are different skills and benefit from different prompts. This is the classic ReAct split applied to code.
  • apply_patch: has to handle diff-application failures cleanly, because the model does generate unapplyable diffs (Section 3).
  • run_tests: needs a sandbox, a timeout, and failure parsing.
  • update_context: the hardest part, how to fold test failures back into the next planning prompt without blowing up the context window.
  • max_iters: the giving-up condition. Without a budget, the agent will loop forever on flaky tests. Claude Code uses 25 as its default ceiling for a single sub-task. Higher budgets help on SWE-Bench scores but bankrupt the agent's cost model in production.

Actor, Critic, or Just Actor?

Some code agent papers use a two-model setup: an actor LLM that proposes edits and a critic LLM that reviews them before they run. OpenHands supports this pattern. In practice, teams have mostly dropped it. Why? Because running tests is a free, perfect critic. The compiler tells you if the code parses; the test suite tells you if it works; the linter tells you if it is clean. Nothing the LLM critic says is as reliable as "the tests pass."

The actor-critic pattern is more useful when feedback is slow, noisy, or absent, which is exactly the situation for research agents (Lesson 56) but not for code. This is one of the few places where code agents are easier than research agents: ground truth is cheap.

Interview Tip

If you find yourself adding an LLM critic to a code agent, first check whether the tests, linter, or type checker can give you the same signal. Free critics beat learned critics whenever they exist, and code has more free critics than almost any other domain.

Real Systems for Context

Concrete anchors help. Here is how the major code agents in 2026 map onto this architecture:

  • Claude Code (Anthropic). Terminal-based. Tools: read, edit, bash, glob, grep, todo. One actor, tight iteration loop, explicit todo list that the agent maintains as state. Aimed at developer workflows, not headless batch runs.
  • Cursor Composer / Agent (Cursor). IDE-integrated. Gets a stream of file context from the editor. Very aggressive context gathering via code graph. Much shorter loop because the human is watching.
  • Devin (Cognition). Long-horizon, autonomous. Runs in a cloud VM with a browser and a shell. Has a scratchpad and an explicit "knowledge" memory store across sessions. Designed for multi-hour tasks a human mostly ignores.
  • SWE-agent (Princeton / academia). Optimized for SWE-Bench. Custom ACI with careful tool design. The paper is the best technical read in this space because the authors document their iteration.
  • Aider. Terminal-based, older, conservative. Maps files to a git repo map, uses unified-diff edits, has a very explicit conversation structure. Its real contribution is the repo map idea, which we will see in Section 2.
  • OpenHands (formerly OpenDevin). Open-source. Supports multiple agent backends. Good for experimenting with the loop itself.

All of these systems share the same core loop. They differ in three places: how they navigate files, how they format edits, and how they handle failure. Those are the three sections that follow.