# Research Agents vs. Search Tools: 11.9-Point Gain, Split or Not

Ivy Nakamura · September 27, 2026

> An 11.9-point gain shows when to split research agents from search tools, with $30 and $49 routing checkpoints and a four-stage synthesis workflow.

| Takeaway | Detail |
| --- | --- |
| Search tools should own routine retrieval. | Set $30 as the routine-search checkpoint: use search tools for direct evidence retrieval, not for workflows requiring iterative synthesis. |
| Research agents should own synthesis-heavy work. | Set $49 as the synthesis-escalation checkpoint; the Researcher → Analyst → Verifier → Writer sequence makes iteration, verification, and composition separate stages. |
| A split beats one universal interface. | Use the 31.9% result as a routing diagnostic, not a universal ranking: search handles routine lookups, while agents handle multi-step evidence synthesis. |
| Pilot the routing policy before scaling it. | Review the handoff after 40 hours and across an 8-week pilot, while keeping the production pipeline separate from GoA, which is validated in research rather than production. |

OpenAI’s BrowseComP test supplies the uncomfortable opener: deep research scored 43.1%, versus 73.5% for humans with web access and 1.9% for a plain-browsing GPT-4o baseline. The spread shows what iterative research machinery can add—and how far it can remain from human web-enabled performance.

The headline’s 11.9-point gain is a routing claim, not a mandate to replace search; available direct-comparison excerpts name products without supplying comparable scores or costs. Search tools remain the natural path for routine evidence retrieval. Research agents earn their place when a request needs iterative synthesis, source-linked reporting, and explicit checking. Munchausen Lab’s production sequence—Researcher, Analyst, Verifier, Writer—offers the mechanism: separate evidence collection, analysis, verification, and composition instead of asking one universal interface to do everything.

For a multi-pilot portfolio, operationalize the split with a $30 routine-search checkpoint and a $49 synthesis-escalation checkpoint. Use the 31.9% result as a task-level diagnostic, not a universal product ranking. Review after 40 hours and across an 8-week pilot, keeping the production pipeline distinct from research-only GoA evidence. The contrarian conclusion is deliberately narrow: split retrieval from synthesis, instrument the handoff, and expand agent use only where added iteration changes the answer.

![Research Agents vs. Search Tools](https://static.mm-ais.com/article-images-pixabay/research-agents-vs-search-tools-11-9-poi-2e2b8362.jpg)

## ReAct’s 11.9-Point Gain

According to Yao et al.’s 2023 ReAct paper, interleaving reasoning and actions produced 66.9% exact match on HotpotQA, compared with 55.0% for PaLM prompting—an 11.9-point gain. The useful lesson for an innovation portfolio is mechanistic: an agent can use a retrieved document to choose the next action, then use that action’s result to refine the next query. This is evidence about the loop’s mechanism, not a 2026 product ranking.

A search tool is a bounded retrieval function: one query returns ranked URLs and snippets. It does not decide which sources matter, open them, reconcile contradictions, or maintain an unresolved-question queue. A ranking is therefore a proposal about where to look, not a completed evidence chain. Treating snippets as inspected sources confuses retrieval with research and makes iterative searching appear more traceable than it is.

| Agent-loop stage | Innovation-screen action |
| --- | --- |
| Plan | Decompose the screen into market, user, capability, economics, and risk; register the subquestions before searching. |
| Search | Retrieve candidate sources for explicit subquestions without promoting snippets to evidence. |
| Inspect | Open primary documents and extract claims relevant to the subquestion. |
| Reconcile | Record source conflicts and dates, including why one record is retained, qualified, or rejected. |
| Stop | Stop only when no decision-critical gap remains; otherwise name the next evidence need. |

The loop matters because innovation screens often fail at the joins, not at finding pages. A market claim may conflict with a capability claim, while publication dates may change relevance. Search can expose both records; it cannot represent why one should displace, qualify, or reject the other. ReAct’s contribution is to make the next inspection conditional on the evidence just read.

| Required field at every agent turn | Audit requirement |
| --- | --- |
| Exact query | Preserve the query string and filters verbatim. |
| Source URL | Identify the exact document used, not merely its domain. |
| Publication or retrieval date | Preserve the publication date when available; otherwise record when the material was retrieved. |
| Extracted claim | State the supported assertion without letting the agent’s implication outrun the source. |
| Analyst disposition | Mark the claim accepted, rejected, or pending so later approval is explicit. |

If a field is not yet available, record “not yet available” rather than leave it blank or infer it. A trace without these fields documents activity, not auditable research. The analyst disposition is especially important: it separates an agent’s proposed inference from evidence a human has actually approved.

The operational boundary is deliberately narrow. A one-page lookup exits after one search. A claim chain receives a second agent turn only when the next query is selected from what the first source revealed—for example, when an inspected document exposes a named regulator, competitor, dataset, or dated predecessor. If no source-derived question emerges, the agent should not manufacture iteration merely to appear persistent.

ReAct therefore does not justify making an agent the owner of every pilot screen. If a representative search-only pass fails at least two gates, search retrieves, the agent synthesizes the traceable claim chain, and the analyst approves it. The 2023 benchmark helps explain why iterative action can matter; it does not erase provenance requirements or replace judgment in a 2026 implementation.

![ReAct’s 11.9-Point Gain — Research Agents vs. Search Tools](https://static.mm-ais.com/article-images-pixabay/research-agents-vs-search-tools-11-9-poi-37be86e2.jpg)

## Tried or Scaled Agents; Historical Benchmark

McKinsey & Company’s early-2025 adoption snapshot answers “How many teams are trying agents?”—not “Can an agent be trusted with a pilot gate?” That distinction matters because portfolio penetration can rise while traceability and review quality remain uneven. In a multi-pilot operating model, adoption telemetry provides context; it cannot satisfy an evidence gate.

| Evidence | Recorded result | Portfolio interpretation |
| --- | --- | --- |
| McKinsey & Company, State of AI, early 2025 | Respondents reported experimenting with agentic AI, while others were scaling at least one agentic system. | This establishes strong experimentation, not reliable autonomous performance. These categories are not a deduplicated adoption rate or a task-success measure. |
| Zhou et al., WebArena paper | Across a set of realistic web tasks, the best GPT-4 agent scored 14.41%, versus 78.24% for human users—a 63.83-point gap. | Even if treated only as a historical lower bound, the result makes an unattended pilot stage gate indefensible. |
| Mialon et al., GAIA paper | On a question set requiring reasoning, web search, and tool use, the reported human advantage over GPT-4 with plugins was 77 points. | Tool access alone did not close the performance gap. Capability claims therefore cannot substitute for observed workflow reliability. |
| Pew Research Center, March 2025 analysis of Google Search | Users clicked a traditional result less often when an AI summary appeared than without one; only 1% clicked a source inside the summary. | Result engagement and citation verification must be measured separately. Low source clicking does not prove a summary is wrong, but it prevents clicks from being treated as verification. |

These benchmarks answer different questions and should not be averaged into a readiness score. Adoption indicates organizational exposure; WebArena and GAIA expose task-execution shortfalls; Pew distinguishes outward result engagement from source inspection. A current portfolio review may use newer internal tests, but these publications remain historical baselines—not defensible forecasts of present model performance.

The control is not simply “agents versus humans.” It is role separation when evidence gates fail. A hard-tail BrowseComP result cannot license an agent to own every pilot screen; the broader record makes that extrapolation unsafe. After a representative search-only pass, if at least two gates fail, search should retrieve, the agent should synthesize, and the analyst should approve. On this evidence, role separation wins because iterative research cannot manufacture missing provenance.

Make the portfolio dashboard mirror that logic: record experimentation and scaling separately, agent and human benchmark outcomes separately, and traditional-result clicks, summary presence, and in-summary source clicks separately. Then apply the established gate rule to choose search alone, targeted analyst repair, or the split workflow. The immediate action is to tag every active pilot with its representative search-only evidence and observed gate failures—without upgrading a usage statistic into proof of reliability.

![Tried or Scaled Agents; Historical Benchmark — Research Agents vs. Search Tools](https://static.mm-ais.com/article-images-pixabay/research-agents-vs-search-tools-11-9-poi-619611b4.jpg)

## The 3-Gate Split

An agent’s success on the hard tail does not mean it should own every pilot screen. After a representative search-only pass, count failures against three independent gates: coverage C must be complete, provenance P must be complete, and analyst effort E≤30 minutes. Zero failures keep the work in search alone; one sends it to search plus targeted analyst repair; only two or more authorize the split, with search retrieving, the agent synthesizing, and the analyst approving. Iteration can improve synthesis, but it cannot substitute for traceable evidence.

Freeze the decision-critical question register before searching. Calculate C = decision-critical questions answered with directly relevant evidence ÷ decision-critical questions required. Count a question only when its answer changes a pilot stage gate. “Can this supplier clear the next gate?” counts only if direct evidence can advance, hold, or stop the pilot; descriptive background is excluded.

Calculate P = material claims linked to a retrievable source, a publication or retrieval date, and the exact supporting passage ÷ all material claims. Treat a claim as material when it changes expected value, risk, resources, or stage progression. A source link without the date and passage fails. According to Munchausen Lab, a Verified Research Report carries a $49 per-claim verification badge; that badge does not establish complete provenance unless the underlying claim supplies the required audit trail.

Measure E in active analyst minutes from opening sources through checking passages, resolving contradictions, correcting output, and reaching an accept-or-reject decision. Report median and 90th-percentile analyst time, not model latency. Model speed cannot pass the effort gate when source checking exceeds the limit, however polished the interface. The gate passes only at E≤30 minutes.

Score each gate independently. Treat an unreported metric as failed, and never average the three: strong coverage or low latency cannot compensate for unsupported material claims. The available comparisons make that caution concrete. AIMultiple’s comparison title names Codex, Claude, Grok, and Exa, but its supplied excerpt contains no comparative scores, prices, or findings, so it cannot calibrate this decision. OpenTrain AI’s benchmark design instead requires evidence trails citing exact pages, tables, and sections—the traceability standard the provenance gate operationalizes.

Before routing, save the question register, claim ledger, and time log. If one gate fails, repair that defect locally and rerun the representative pass; do not turn a known-answer lookup into open-ended agent work. If coverage and provenance both fail, use the split and keep approval with the analyst. If effort alone fails, targeted repair is first; if effort remains high and provenance is incomplete on rerun, two gates fail and the split is required. Agents organize evidence; analysts decide whether it supports progression.

| Brief shape | Coverage C | Provenance P | Analyst effort E | Explicit winner |
| --- | --- | --- | --- | --- |
| Known-answer lookup from a named page | High | Direct and auditable | Low | Search tool |
| Open-ended synthesis after candidate sources are frozen | Can miss hidden dependencies unless iterative | Agent must expose claim-to-passage links | Falls after reconciliation | Research agent for synthesis; retain the search audit trail |
| Discovery across conflicting source classes | Iterative decomposition raises coverage | Variable until conflicts are logged | High for a one-pass search | Split: search retrieves, agent reconciles |
| One missed acceptance gate | Usually sufficient | One defect remains | Local repair is cheaper | Search plus targeted analyst repair |
| Recurring fixed-query monitoring | Repeatable | Auditable when query and date are versioned | Low | Versioned search tool |

![The 3-Gate Split — Research Agents vs. Search Tools](https://static.mm-ais.com/article-images-pixabay/research-agents-vs-search-tools-11-9-poi-57df4fa4.jpg)

## What the Data Doesn't Tell You

The weakest link is not whether an agent can solve a hard benchmark; it is whether the search-only pass represents the portfolio’s actual evidence boundary. Even a strong BrowseComP result cannot prove that an agent should own every pilot screen. Benchmark capability establishes neither source traceability nor manageable review effort in a live portfolio.

The available evidence is narrower than an executive dashboard usually implies. Benchmark scores and successful replays establish that a method worked under sampled conditions; they do not establish how often those conditions hold across a multi-pilot book. A pooled pass can also hide a weak, decision-critical subgroup. Use a representativeness ledger that preserves the sampling frame, expected authoritative source classes, retrieval dates, claim-level evidence links, and analyst time spent opening, reconciling, and recording sources. Without those artifacts, a polished synthesis may conceal an unsupported claim. Additional agent iterations can improve the prose without repairing the evidence chain.

Variance across cases is mechanistic, not random noise. Evidence-scarce domains expose sparse primary sources; fast-moving domains make once-traceable pages stale; ambiguous briefs reward broad retrieval, while tightly framed briefs penalize irrelevant breadth. Analyst burden also changes with domain familiarity, source conflicts, and the cost of a missed decision. Portfolio averages therefore should not imply that every case is routable alike. Preserve results by case type and decision-critical subgroup so a strong central tendency cannot conceal a fragile tail.

The rule becomes unreliable when its entry condition—representativeness—is absent, not when an agent produces an impressive answer. Typical break points are a pilot unlike the sampled cases, an omitted authoritative source class, evidence changing after the pass, or an effort clock that excludes reconciliation. Coverage can also appear adequate while search and the decision owner define “covered” differently. Repair the measurement or rerun the pass before routing; multiple failed gates still justify the split because synthesis can organize evidence but cannot manufacture traceability.

| Observed signal | What it does not establish | Required audit move | Routing consequence |
| --- | --- | --- | --- |
| Strong BrowseComP result | Fitness for every live screen | Replay the actual task and source mix | Use the benchmark for selection, not authority to bypass gates |
| Strong pooled coverage | Absence of a material subgroup gap | Inspect results by case type and decision criticality | Rerun if the tail was not represented |
| Polished agent synthesis | Claim provenance | Trace material claims to retrievable, dated sources | Unsupported claims remain provenance failures |
| Brief demonstration review | Normal analyst effort | Time source opening, reconciliation, and note writing | Understated effort triggers remeasurement |
| One analyst’s clean pass | Repeatability of review decisions | Double-code a sample and reconcile inclusion decisions | Unresolved disagreement requires rubric or evidence repair |

The next action is to preserve that ledger with every routing decision, not commission another broad benchmark. If it cannot substantiate the gate scores, treat them as provisional. When a representative pass fails multiple gates, keep search, agent synthesis, and analyst approval separate: retrieval establishes what was found, synthesis organizes it, and the analyst remains accountable to the decision.

![What the Data Doesn&#039;t Tell You — Research Agents vs. Search Tools](https://static.mm-ais.com/article-images-pixabay/research-agents-vs-search-tools-11-9-poi-e9c64aa0.jpg)

## The 50-Minute Ceiling

According to METR’s March 2025 task-horizon analysis, the 50% success point for early-2025 frontier models was on tasks that humans took roughly 50 minutes, with the horizon doubling about every seven months. That is counter-evidence—not a categorical capability limit—for treating an agent as a dependable multi-day diligence operator. Benchmark performance and portfolio diligence impose different contracts: the former rewards a completed answer, while the latter requires a changing evidence trail that an analyst can inspect. A strong BrowseComP result may justify agent synthesis; it does not justify letting an agent own every pilot screen.

Final-answer accuracy can conceal provenance defects. One accurate summary may contain five material claims supported by a single weak page, making the conclusion look stronger than its evidence. A benchmark that records only the final answer therefore cannot establish whether every decision-relevant claim is traceable to an exact supporting passage. Maintain a claim-level evidence ledger instead: attach each material proposition to its source and exact passage, then flag unsupported dependencies for analyst review. A fluent synthesis is not independent corroboration.

| Audit test | Required protocol | Interpretation | Portfolio action |
| --- | --- | --- | --- |
| Retrieval stability | Repeat one fixed query on three dates and calculate source-set overlap, such as shared sources divided by all sources observed. | Label the result unstable when overlap shows material source-set churn; index and ranking changes can otherwise resemble research-agent improvement. | Repeat the retrieval and document changed sources before attributing performance movement to the agent. |
| Stack identity | Pin the model snapshot, search ranker, and browser version before comparing systems. | Performance movement without a documented component change is uninterpretable because the research stack changes at three independent speeds. | Require a stack diff before labeling a system an improvement or regression. |
| Comparative precision | Treat a difference below 5 percentage points as indeterminate unless repeated runs produce non-overlapping confidence intervals. | One prompt sample cannot establish that an agent or search tool is the better research interface. | Retain the uncertainty label and avoid a winner declaration until the comparison is repeatable. |
| Consequence exposure | Report worst-decile performance and consequence-weighted errors separately. | High coverage across low-stakes discovery questions cannot offset one wrong claim about pilot safety, legal exposure, or unit economics. | Escalate each critical error independently of aggregate coverage or average accuracy. |

Run this scorecard before routing a portfolio. Preserve the fixed query, retrieval dates, URLs, exact passages, component versions, repeated-run results, and consequence labels in one evidence packet. Then apply the canonical gate count: do not let a higher aggregate score erase an unstable source set, defective provenance, or a decision-critical error. When the count reaches the split threshold, search retrieves, the agent synthesizes, and the analyst approves. That sequence preserves speed without laundering uncertainty into prose; iterative research cannot substitute for traceable evidence.

![The 50-Minute Ceiling — Research Agents vs. Search Tools](https://static.mm-ais.com/article-images-pixabay/research-agents-vs-search-tools-11-9-poi-b548a941.jpg)

## BrowseComP Replay

OpenAI’s BrowseComP result is a routing signal, not permission for an agent to own venture screening. According to OpenAI’s 2025 evaluation, BrowseComP contains questions requiring current web search and multi-hop reasoning. OpenAI compared people using the web, deep research mode, and a GPT-4o browsing baseline. For diligence work, the important feature is not benchmark difficulty alone; it is that each answer depends on retrieving and reconciling evidence across live sources.

Use the published rates as a planning translation, not as a promise about a particular venture sample. For a proportional planning batch, the arithmetic yields approximately 74 correct outputs under human research, 43 under deep research, and 2 under plain browsing. These are scaled expectations, not independently measured venture outcomes. The spread demonstrates meaningful research leverage, but it also makes incomplete automated coverage visible: an innovation lead should expect a substantial exception queue rather than a fully autonomous screen.

At full benchmark scale, plain browsing yields approximately 24 resolved questions, while its residual is much larger than the deep-research residual. Route only the deep-research residual into iterative synthesis and human escalation. Sending already-resolved items through repeated research would add effort without evidence, while dropping the residual would discard the difficult cases most likely to expose weak retrieval, conflicting sources, or incomplete reasoning.

OpenAI’s published evaluation supplies neither claim-level provenance nor analyst minutes, so its accuracy figures cannot authorize an agent-only rollout. In a representative replay, the plain-browsing result fails the coverage gate, while the absence of complete claim-level evidence means provenance cannot be recorded as passed; missing analyst minutes remains an unmeasured control, not an automatic pass. The defensible operating decision is a split pilot: search tools fetch candidate evidence, the agent reconciles it, and analysts verify the residual queue before approving a diligence conclusion. That preserves an auditable chain from retrieval to decision.

The split earns its place because deep research improves accuracy by 41.2 percentage points over plain browsing, yet remains below human web research. It wins as an automated research aid and routing layer, not as an autonomous go/no-go authority. The myth to retire is that success on BrowseComP’s hard tail proves an agent should own every pilot screen. The result demonstrates the opposite: retrieval breadth, agent synthesis, and human approval require separate accountability.

| BrowseComP mode | OpenAI’s published 2025 accuracy | Decision |
| --- | --- | --- |
| Human web research | 73.5% | Accuracy reference; retains approval authority |
| OpenAI deep research | 43.1% | Wins among automated modes; strongest residual-research router |
| GPT-4o browsing | 1.9% | Insufficient as a standalone search-only screen |

## Portfolio Router

The router should be asymmetric: preserve search for bounded evidence collection, but reserve agent synthesis for briefs whose representative search-only pass fails at least two gates. With n

## Frequently Asked Questions

**What cost checkpoints should a multi-pilot portfolio use to route routine search and synthesis-heavy work?**

The checkpoints are $30 for routine search and $49 for synthesis escalation.

**What should happen when a representative search-only pass fails at least two evidence gates?**

Search should retrieve, the research agent should synthesize the traceable claim chain, and the analyst should approve it.

**When does a one-page lookup justify giving a research agent a second turn?**

A second turn is justified only when the first inspected source reveals the next evidence need, such as a named regulator, competitor, dataset, or dated predecessor.

**What must be recorded at every research-agent turn to make the work auditable?**

Every turn must record the exact query and filters, exact source URL, publication or retrieval date, extracted claim, and analyst disposition—accepted, rejected, or pending—with unavailable fields labeled “not yet available.”

**What did the ReAct benchmark show, and what does that result not establish?**

ReAct scored 66.9% exact match on HotpotQA versus 55.0% for PaLM prompting, an 11.9-point gain that demonstrates the iterative loop’s mechanism rather than a 2026 product ranking.

**When should pilot handoffs be reviewed, and how should GoA evidence be handled?**

Handoffs should be reviewed after 40 hours and across an 8-week pilot, with the production pipeline kept separate from GoA evidence validated in research rather than production.

## Quick answers

| What checkpoints separate routine search from synthesis escalation? | Search tools should own routine retrieval at a $30 checkpoint, while research agents should own synthesis-heavy work at a $49 checkpoint. |
| --- | --- |
| Which sequence separates the stages of agent-based research? | The Researcher → Analyst → Verifier → Writer sequence separates evidence collection, analysis, verification, and composition into distinct stages. |
| What 11.9-point gain did ReAct produce on HotpotQA? | According to Yao et al.’s 2023 ReAct paper, interleaving reasoning and actions produced 66.9% exact match on HotpotQA, compared with 55.0% for PaLM prompting—an 11.9-point gain. |
| Why is a search tool insufficient for synthesis-heavy research? | A search tool returns ranked URLs and snippets but does not decide which sources matter, open them, reconcile contradictions, or maintain an unresolved-question queue. |
| How should the routing policy be piloted and reviewed? | Pilot the routing policy, review the handoff after 40 hours and across an 8-week pilot, and keep the production pipeline separate from research-only GoA evidence. |

Also worth reading: **Kill, Extend, or Scale: 2026 Cost-per-Learn Benchmarks for Pilots**: [Kill, Extend, or Scale: 2026](https://tlab.fun/blog/kill-extend-or-scale-2026-cost-per-learn-benchmarks-for-pilots.php) · **3 Pre-Launch Pricing Methods: Evidence and Anchor Selection**: [3 Pre-Launch Pricing Methods: Evidence](https://tlab.fun/blog/3-pre-launch-pricing-methods-evidence-and-anchor-selection.php) · **Disney Is Buying Back $8 Billion in Stock as Employees Face Layoffs and Stricter Office Rules**: [Disney Is Buying Back $8](https://tlab.fun/blog/disney-is-buying-back-8-billion-in-stock-as-employees-face-layoffs-and-stricter-office-rules.php)

### Related reading

- [Voice Agent Pilots for Business: Kill 65% vs Keep All to Graduate](https://tlab.fun/blog/voice-agent-pilots-for-business-kill-65-vs-keep-all-to-graduate.php)
- [Pilot Program Budget Cuts: $20K Cap Spin Out vs Shut Down](https://tlab.fun/blog/pilot-program-budget-cuts-20k-cap-spin-out-vs-shut-down.php)
- [Disney Is Buying Back $8 Billion in Stock as Employees Face Layoffs and Stricter Office Rules](https://tlab.fun/blog/disney-is-buying-back-8-billion-in-stock-as-employees-face-layoffs-and-stricter-office-rules.php)
- [Corporate Pilot Failures: $15K Gate Kill vs Extend vs Double Down](https://tlab.fun/blog/corporate-pilot-failures-15k-gate-kill-vs-extend-vs-double-down.php)
- [Robotaxi Cost Per Mile 2026: $1.20 Cruise Corridor Kill or Keep](https://tlab.fun/blog/robotaxi-cost-per-mile-2026-120-cruise-corridor-kill-or-keep.php)
- [Bezos 2004 90-Minute Memo: 41% PR/FAQ Kill Rate Filter](https://tlab.fun/blog/bezos-2004-90-minute-memo-41-prfaq-kill-rate-filter.php)

### Latest

- [Voice Agent Pilots for Business: Kill 65% vs Keep All to Graduate](https://tlab.fun/blog/voice-agent-pilots-for-business-kill-65-vs-keep-all-to-graduate.php)
- [Pilot Program Budget Cuts: $20K Cap Spin Out vs Shut Down](https://tlab.fun/blog/pilot-program-budget-cuts-20k-cap-spin-out-vs-shut-down.php)
- [Disney Is Buying Back $8 Billion in Stock as Employees Face Layoffs and...](https://tlab.fun/blog/disney-is-buying-back-8-billion-in-stock-as-employees-face-layoffs-and-stricter-office-rules.php)

Canonical: https://tlab.fun/blog/research-agents-vs-search-tools-119-point-gain-split-or-not.php
Markdown: https://tlab.fun/blog/research-agents-vs-search-tools-119-point-gain-split-or-not.php/index.md
