Why Autonomous Research Agents Hallucinate — and How a Critic Loop Surfaces Unsupported Claims
Research agents can look authoritative and still invent citations. A critic with source access catches unsupported claims — it doesn't erase risk.
Why this matters
An autonomous research agent that hallucinates can be worse than no agent at all. It can produce authoritative-looking outputs with plausible citations that are partly or entirely wrong. The researcher receiving the output has no efficient way to verify it, which is often why the work was delegated to an agent in the first place.
A critic is a second pass that must cite sources or flag gaps, not another confident writer.
I've built research agents that failed quietly in exactly this way. The planner generated useful sub-questions. The executor retrieved real documents. Then, under context pressure, the synthesis step tried to turn partial results into a coherent report. It filled gaps with claims that had no source.
The claim was not in any retrieved document. It came from the model's training data and appeared in the report as though it had been researched.
The failure is architectural, not just a prompt-engineering issue. A planner-executor pair without a verification step cannot reliably distinguish between "found in a source" and "the model is confident about this." Adding a critic with independent source access is a structural mitigation. Here is why, and what it looks like in practice.
1. The planner-executor pattern and its limits
The planner-executor pair is a common architecture for autonomous research. A planner decomposes a high-level topic into sub-questions. An executor retrieves and summarizes answers to each sub-question. The planner then synthesizes the executor's outputs into a final report.
This works when the corpus is well-defined, the sub-questions can be answered independently, and the retrieved documents contain enough relevant information. It breaks down when the corpus is sparse, the sub-questions are ambiguous, or the synthesis step must reconcile incomplete and contradictory sources.
The synthesis failure mode is the dangerous one. The planner receives executor outputs such as "Source A says X" and "Source B is ambiguous about Y." It still has to produce a coherent report. With limited tokens and many sources, the model may fill ambiguous sections with the answer it expects, based on its training data.
The result sounds coherent. The claims are not verifiable.
I do not expect a critic loop to fix sparse corpora or ambiguous questions. Its role is narrower: to surface the failure. It flags claims that cannot be mapped to retrieved sources and gives the planner enough signal to re-query or explicitly mark uncertainty.
2. What a critic agent actually does
The critic is not a second synthesizer. I use it as a verification agent with two specific capabilities: access to the same source documents the executor retrieved, and a strict grounding constraint.
The task is intentionally narrow. Given the executor's summary and the source documents, the critic marks every claim as grounded (found verbatim or paraphrased in a source), inferred (logically follows from sources but is not stated directly), or unsupported (not found in any retrieved source).
from pydantic import BaseModel
from typing import Literal
class ClaimVerification(BaseModel):
claim: str
status: Literal["grounded", "inferred", "unsupported"]
source_url: str | None # require it in application validation when grounded
confidence: float # validate this is between 0.0 and 1.0
class CriticOutput(BaseModel):
verified_claims: list[ClaimVerification]
overall_groundedness: float # fraction of claims that are "grounded"
flags: list[str] # specific issues for the planner to act on
I keep this output structured because the planner needs a stable control signal. The critic does not return free text, and it does not route its result back to the executor.
The planner reads overall_groundedness and the flags list, then selects the next step. If groundedness exceeds a configured threshold, such as 0.8 in a system where that threshold has been calibrated against task risk, the report can be approved. If it falls below that threshold, the planner re-queues flagged sub-questions with explicit instructions to find sources for unsupported claims.
The loop also needs a max-iterations guard. In my own implementations, I prefer a hard stop to an agent that keeps searching until it finds something plausible. If groundedness does not improve after a small fixed number of re-query cycles, the planner explicitly marks low-confidence claims in the final report rather than silently passing them through.
3. Independent source access for the critic
The critic must have access to the same source documents the executor retrieved. It cannot rely only on the executor's summary of those documents. I do not relax this requirement.
If the critic sees only the executor's summary, it can check whether the summary contradicts itself. It cannot verify whether a claim actually appears in the source. A summary can be internally consistent and still be wrong.
# Wrong: critic sees only the summary
critic_input = {
"summary": executor_output.summary,
"task": "Verify the claims in this summary."
}
# Critic can check consistency but not source grounding
# Right: critic sees summary and original documents
critic_input = {
"summary": executor_output.summary,
"source_documents": executor_output.retrieved_docs, # original text, not summaries
"task": "For each claim in the summary, verify it against the source documents."
}
# Critic can map claims to specific passages
The practical constraint is context size and retrieval cost. The critic's context window must contain the summary and relevant source excerpts, or it must be paired with a retrieval step that fetches those excerpts deterministically.
For long-form research tasks with many sources, I usually use one of two approaches. I either run the critic one sub-question at a time, or I retrieve source passages specifically for the critic pass.
Running per sub-question keeps cost under control, but it can miss multi-source claims such as "Sources A and B both confirm that..." unless the full claim context is passed explicitly. Running the critic against a broader corpus costs more, but it can catch cross-source contradictions. For example, it can identify sources that provide different dates, revenue figures, or regulatory interpretations for the same claim.
4. Recursive summarization and citation tracing
Autonomous research agents often retrieve more content than fits in a single context window. The standard response is recursive summarization: summarize document A, summarize document B, then synthesize the summaries.
Recursive summarization is useful for compression. It is risky for citation tracing.
When document A is summarized, the condensed version can lose specific passages. The critic can no longer map claims to the exact text in document A. The evidentiary chain is broken.
A simple way to preserve it is to store the original document alongside the summary, then pass both to the critic or make the original retrievable by document ID and span.
class ExecutorResult(BaseModel):
sub_question: str
summary: str # compressed, used for planning
source_docs: list[str] # original text, used for critic verification
source_urls: list[str]
The critic reads the summary to understand claim context. It then searches the original text for the specific passage.
This increases the data stored for each executor result, which is acceptable when the corpus is small enough to retain directly. If corpus size makes that infeasible, I would embed or index the originals and run citation search at critic time instead of passing them inline. The important part is durable state: the system must preserve enough original source material to support deterministic recovery of the evidence chain.
5. Map-reduce for independent sub-questions
When a research task decomposes into independent sub-questions — market trends, competitive landscape, regulatory environment — the planner can dispatch the executor in parallel (map) and aggregate verified results (reduce).
The map step dispatches all sub-questions simultaneously. The reduce step synthesizes verified executor outputs into a final report. Depending on context size and cost constraints, the critic can run on each mapped result, on the final synthesis, or on both.
A final critic pass keeps verification attached to the final output. If the reduce step changes wording, merges claims, or introduces cross-source conclusions, those claims still need to be checked against the retrieved corpus.
The failure mode to avoid in the reduce step is cross-question drift. The planner may synthesize across sub-questions without recognizing that the same claim appears in multiple executor results with conflicting values. "The market size is $5B" from one sub-question and "The market size is $3B" from another should be flagged, not averaged. The critic's flags list surfaces these conflicts.
6. Handling coverage gaps explicitly
A research agent working with sparse sources needs explicit coverage-gap handling. The critic loop must also handle cases where retrieved sources are insufficient to answer the question.
The signal is repeated unsupported claims. If the critic consistently marks claims as unsupported across re-query cycles with different search terms, the problem may be corpus coverage rather than query quality. In that case, I prefer to report the coverage gap explicitly instead of producing a plausible-sounding answer.
class ResearchReport(BaseModel):
findings: list[ClaimVerification]
coverage_gaps: list[str] # topics where sources were insufficient
confidence: float # overall groundedness across all findings
generated_at: str # ISO timestamp
The coverage_gaps field is required. I make the planner fill it, even when the value is an empty list. A report schema without a required coverage-gap field can allow gaps to be omitted.
Explicit coverage gaps are more useful than confidently wrong answers. A user who sees "Coverage gap: regulatory landscape in EU after 2023 — sources available only through Q2 2023" knows what additional research is needed. A user who receives a hallucinated answer has no signal that anything is wrong until downstream validation fails.
The category frame
Autonomous research agents that produce verifiable outputs need three components: a planner that decomposes and synthesizes, an executor that retrieves and summarizes while preserving sources, and a critic that checks claims against original documents.
The critic loop does not eliminate grounding errors. It makes unsupported claims visible and actionable. The planner can then preserve the distinction between sourced claims, justified inferences, and coverage gaps rather than flattening them into false confidence.
The goal is for every claim to be traceable to a source URL, explicitly marked as inferred, or surfaced as a coverage gap. The system should not silently create a fourth category.
For this to remain reliable over time, I also need reproducible evals: fixed tasks, frozen source snapshots where possible, expected claim classifications, and regression checks for unsupported claims that previously slipped through. Without that layer, a critic loop can look correct in a demo and drift as prompts, models, retrieval settings, or source content change.
If you're building research automation for compliance, legal, or financial workflows, write me and I can review the verification architecture against the checks described here. I can help examine where source preservation, critic access, coverage-gap reporting, and evaluation fit into the system design.
FAQ
Why does a planner-executor research agent hallucinate during synthesis?
I see the dangerous failure in the synthesis step. The planner receives partial or ambiguous executor outputs and must produce a coherent report under context pressure. In that setting, the model may fill gaps with claims from training data. Those claims can then appear as though they were found in retrieved sources.
What does a critic agent verify in an autonomous research workflow?
I use the critic as a narrow verification agent, not a second synthesizer. It receives the executor summary and source documents. For each claim, it returns a status: grounded, inferred, or unsupported. It also returns structured output with planner-facing flags and an overall groundedness score.
Why must the critic see the original source documents?
I do not let the critic rely only on the executor summary. A summary can be internally consistent and still be wrong. The critic needs the original retrieved documents to map claims to specific source passages.
How should recursive summarization preserve citation tracing?
I keep the original document alongside the compressed summary, or I store it in an index that can retrieve the relevant span later. The summary is useful for planning, but it can lose the exact passages needed for verification. The critic reads the summary for claim context, then checks the original text for the supporting passage.
What should happen when claims remain unsupported after re-querying?
I treat repeated unsupported claims across re-query cycles with different search terms as a likely coverage gap, not merely a query failure. The planner should report the gap explicitly. It should not pass through a plausible-sounding answer with false confidence.
Related articles
Emotional Memory
I check this article against the public proof at DOI 10.5281/zenodo.19972258.
Oct 4, 20261 min read#Memory#LLM#Research#StateRAG in Production: Fix Chunking and Re-Ranking Before Touching Embeddings
When retrieval is weak, swapping embeddings rarely fixes it. Diagnose chunking and re-ranking first.
Dec 20, 202412 min read#RAG#Retrieval#LLM#ProductionFive Function-Calling Patterns I Use in Production
Five function-calling patterns I use in production systems, and the anti-patterns they replaced.
Nov 25, 202413 min read#LLM#Function Calling#OpenAI#Production