0%
Applied AI Systems
Retrieval and Knowledge
Agents and Reasoning
Evaluation, Interpretability, and Safety
Production Research Engineering
Building a Research Agent Case Study
Perplexity answers research questions with cited sources in about five seconds. Gemini Deep Research spends ten minutes producing a twenty-page report with hundreds of references. Grok DeepSearch streams its thinking live while browsing the web. You.com, Phind, GPT Search, and Anthropic's Claude with web search are variations on the same theme. They look magical, but under the hood they are all assembling the same small set of pieces in slightly different ways.
This lesson is a case study. We are going to reverse-engineer a production research agent end to end, from the user's question to the cited answer, and leave you with a system you could actually build. This is not a survey. We are going to commit to specific design choices and defend them.
What a Research Agent Must Do
A research agent takes a natural-language question and returns a written answer grounded in web sources. That sentence hides a surprising amount of complexity. Let us unpack what the system has to handle.
First, questions come in very different shapes. Some are narrow and factual ("What was the market cap of Nvidia in Q4 2025?"). Some are broad and comparative ("How do modern reasoning models differ from standard LLMs?"). Some are multi-hop ("Which of the top ten AI labs by revenue has the largest compute budget per researcher?"). A good research agent has to handle all of these without a different code path for each.
Second, the open web is noisy. Search engines return pages that range from authoritative (arXiv, company blogs, major news outlets) to garbage (SEO farms, stale Wikipedia mirrors, LLM-generated slop). The agent has to filter and rank sources, not just fetch whatever the search API returns first.
Third, the user expects citations. Every non-trivial claim in the answer should trace back to a source the user can click and verify. This is the single biggest thing that separates a research agent from a chatbot. When Perplexity says "Nvidia's revenue grew 265% year over year", the number is clickable and goes to the exact paragraph in the source that stated it. Without that trail of evidence, users cannot trust the output, and they should not.
Fourth, the answer has to be fast enough to be usable. For an interactive question, that means five to fifteen seconds end to end. For a deep research task, it means streaming partial results while the full report is still being produced. The architecture has to support streaming.
The Minimum Tool Set
Every research agent I have studied uses roughly the same six building blocks. In order of the pipeline:
- Query understanding: an LLM call that takes the user's question and produces a plan, a set of sub-questions or search queries, and any clarifying rewrites.
- Web search: a search API that returns a ranked list of URLs with titles and short snippets. Tavily and SerpAPI are the two most common choices for agent builders. Perplexity, You.com, and GPT Search all use a mix of their own crawl and partnerships with Bing or Google.
- Page fetch and scrape: given a URL, fetch the page and extract the readable content. This is harder than it sounds because half the web is behind JavaScript, paywalls, or bot detection. Production systems use a headless browser, a reader-mode extractor like Mozilla Readability or trafilatura, and a cache.
- Reranker and chunk selector: not every paragraph of a 5000-word article is relevant. A reranker (often a cross-encoder like Cohere Rerank or a distilled BGE model) scores candidate chunks against the query and keeps only the top few per source.
- Synthesis LLM: the big model call that takes the question, the top chunks from the top sources, and produces the written answer with inline citations.
- Citation verifier: a final step that checks whether each claim in the answer is actually supported by the cited source, and either fixes or flags claims that are not.
That is it. Six components, maybe 200 lines of glue code for a basic version. Production systems add streaming, caching, query routing, follow-up question handling, and a dozen other niceties, but the core is this.
The Research Agent Loop in Pseudocode
Here is the whole thing as code. Read this, then we will spend the rest of the lesson unpacking each piece.
Notice what is not here. There is no ReAct-style think-act-observe loop. There is no planner that decides the next step. There is no memory of past interactions. This is deliberate. For the vast majority of research questions, a straight-through pipeline works better than a free-form agent loop, because the failure modes are easier to reason about and the cost is bounded.
Anthropic's Building Effective Agents blog post makes this point explicitly: workflows (fixed pipelines with LLM steps) are more reliable and cheaper than agents (LLMs making decisions about what to do next), and you should reach for a workflow first. A research agent is primarily a workflow. Only when the workflow fails (the synthesis pronounces the question unanswerable, the verifier rejects the claims) should you fall back to something more interactive.
Why Not Just Use a Vector Database?
A reader who has taken the RAG lessons might ask: why not just embed the web into a vector database and do semantic search? This is what ChatGPT did with Bing integration before GPT Search existed.
The answer is that the web is too big and too fresh. Perplexity crawls and indexes tens of millions of pages per day. The delay between a news story being published and it being retrievable has to be under a few minutes for the agent to feel current. Traditional web search engines have spent twenty years optimizing this exact problem, so research agents mostly use existing search APIs rather than reinventing the index.
The vector database approach works well for closed corpora: your company's documentation, a specific book, a dataset of scientific papers. For the open web, keyword search plus LLM-powered query rewriting plus reranking beats pure semantic search on most metrics. Perplexity uses a hybrid: search APIs for recall, semantic reranking for precision. We will follow the same pattern.
The single biggest architectural decision for a research agent is whether it is a workflow or a free-form agent. Workflows are pipelines with LLM steps, predictable, cheap, and testable. Agents decide their own next action at each step, flexible, expensive, and hard to debug. Start with a workflow. Only reach for agents when the workflow genuinely cannot handle the task. Perplexity, You.com, and Gemini Deep Research are all primarily workflows with a small amount of branching.
Latency and Cost Budget
Before we dive into components, it helps to know what we are spending. A representative research query for an interactive agent might spend:
- Query understanding LLM call: 300ms, 1000 input tokens, 200 output tokens
- Web search: 400ms, 3-5 parallel calls to the search API
- Page fetch: 1500ms, fetching 5-10 URLs in parallel (the slowest step)
- Reranking: 300ms, one batch call to a reranker API
- Synthesis LLM call: 2500ms, 8000 input tokens, 600 output tokens
- Citation verifier: 800ms, one or two additional LLM calls
Total wall-clock: roughly 6 seconds. Total LLM cost at gpt-4o-mini-ish rates: roughly 2 cents per query. Search API cost: another half cent. Page fetch cost: effectively free if you use a shared headless browser pool.
The big knobs are how many search queries you issue, how many pages you fetch, and which model you use for synthesis. A premium research tier (Perplexity Pro, Gemini Deep Research) bumps every number up by 5-10x: more queries, more pages, a bigger synthesis model, and additional verification passes.
Deep research modes are even more expensive. Gemini Deep Research can spend several dollars per report and take ten minutes of wall-clock time, because it is running 50-100 searches and reading hundreds of pages. The architecture is the same as the interactive version, just with higher budgets on every step.
What This Lesson Is and Is Not
We are going to build a concrete interactive research agent, the kind that returns an answer in five to ten seconds. The techniques generalize to deep research, but we will not spend much time on the orchestration of hundred-step workflows. That is a separate engineering problem.
We are also going to lean heavily on Perplexity as the reference system, because it is the most studied research agent in production and the company has been relatively open about their architecture. Where other systems (You.com, GPT Search, Gemini Deep Research, Grok DeepSearch) make different choices, I will call them out.