Source Evaluation and Credibility Scoring in Research Agents
Separate source credibility from overall quality scores to catch citation failures before they ship.

Three weeks after a research agent ships its brief, someone finally clicks citation 7. The paper it points to says the opposite of what the brief claimed. Nothing in the dashboard caught this. The aggregate quality score passed. The writing read smoothly. The latency line stayed green the whole time. That gap between a clean-looking run and a broken citation is the subject of this piece, and it comes down to one design flaw: credibility evaluation gets folded into one overall quality score instead of being measured on its own.
A single number can't tell you which stage failed, and that's the trap. A brief can score close to perfect on citation validity and still be worthless, because every source it cites traces back to the same upstream paper, dressed up as five independent sources. A brief can also be fully grounded in solid evidence and still miss the point, because it answers a question next door to the one that was asked. Averaging these failure types into one score doesn't cancel them out; it just hides them from each other.
Part of why this hides so well comes down to how these models write. Any decent LLM produces prose that sounds confident and well-sourced, a fluency that has nothing to do with whether the sourcing underneath actually holds up. Fluent writing is not evidence of fluent reasoning, and the surface metrics most teams track, like read time or grammar, can't tell the difference.
The way the field currently measures these systems makes the problem worse rather than better. Most evaluation of LLMs that cite web sources checks whether the final answer is correct rather than whether the evidence behind it is any good. That's measuring the wrong thing: a lucky guess backed by garbage sources scores the same as a well-reasoned answer backed by garbage sources, and neither gets flagged for the sourcing problem. A survey of autonomous research agents from Carnegie Mellon and collaborators calls this the "verification gap": releasing code with a system is now standard practice, but only 38% of the systems surveyed report any method for verifying novelty, so the checks a reviewer would need to confirm a result actually holds up simply aren't there for most systems.
The numbers on citation accuracy make the stakes concrete. A cross-model audit spanning 40 research domains found citation hallucination rates that swung from low to alarmingly high depending on the model and the domain, with fabricated citations in computer science research climbing to extreme levels for some models. Even commercial deep research tools built specifically for this job aren't immune: citation accuracy errors run from the single digits up past twenty percent, which is a real problem when a business decision is riding on what that citation actually says.
What a five-stage pipeline looks like
Every deep-research agent, no matter who builds it, runs through the same five-stage process. Source credibility evaluation is stage three, sitting right in the middle, and a failure there poisons everything that comes after it. The table below lays out each stage, what it's responsible for, and the specific ways it tends to break.
| Stage | Job | Common failure modes | |---|---|---| | 1. Query planning | Break the question into sub-questions and allocate research effort | Scope drift, a missing sub-question, too much effort spent on one angle | | 2. Source retrieval | Pull candidate sources from the web | Stopping too early, recursive loops that never resolve, gaps in the retrieval log | | 3. Source evaluation | Judge whether each retrieved source is trustworthy for the question | Using a low-tier source to answer a high-stakes question, relying on a monoculture of sources that all say the same thing | | 4. Claim-evidence alignment | Match specific claims in the brief to specific evidence | Citing the wrong passage, paraphrasing past the point of accuracy, inventing a citation outright | | 5. Synthesis coherence | Turn aligned claims into a coherent final brief | Stitching paragraphs together without a throughline, answering a nearby question instead of the one asked, ignoring the requested format |
Treat these five as a vector rather than a single score. Average them together and a collapse in one stage barely moves the overall number, and worse, nothing in that number tells anyone which stage to go fix.
Stage three is a chokepoint in this pipeline. Whatever clears source evaluation becomes the evidence that stages four and five build on, so a bad source that slips through gets cited as grounding in the final brief with no flag anywhere that it shouldn't have been there. Most teams that track pipeline quality at all have dashboards with nothing watching the plan stage or the synthesis stage. Source evaluation often gets the same treatment, left unmonitored because everyone assumes the retriever already handed over good material.
Why source credibility is harder to score than it looks
Credibility is a judgment made across multiple independent dimensions, and when a team collapses all of those into one quality label, the specific thing that's wrong with a bad source disappears along with the label's usefulness.
SourceBench, a benchmark published under arXiv number 2602.16942, is the most developed attempt so far at scoring web source quality for AI answers, and it breaks the old habit of scoring sources purely on relevance to the query. Instead it scores eight dimensions split across two categories: content quality and page-level signal. The content side covers how relevant a source is to the query, how accurate its claims are, and how neutral or objective it reads. The page-level side covers how current the information is, how clear the organization and layout are, and how much authority the author or domain carries.
A retrieved source's value to a person doesn't come down to how closely its text matches the search query. It comes down to the whole experience of actually trying to use it. A page that gets every fact right but buries them under a wall of intrusive ads, or never states when it was published, adds friction that breaks trust in exactly the high-stakes situations where trust matters most. SourceBench tested this across search-equipped LLMs, traditional search engine results, and AI-native search tools, and found real gaps in source quality between these system types. Its automated scoring also tracked closely with human judgment, staying within a mean error of under 0.5 across every metric tested.
The broader research on how people judge credibility backs up the same point from a different angle. A University of Salerno literature review pulls together what drives credibility judgments, and the list includes the quality of the site itself, how the page is structured, whether bias shows up in the content, how professional the authors seem, and whether the source has a track record of spreading misinformation. No single factor on that list decides the verdict alone. The same review makes a point that static, one-time scoring tends to miss: the line between reliable and unreliable isn't fixed in place. A source crosses it gradually, as problems pile up over time, so a credibility score taken once and never revisited will eventually be scoring a source that no longer exists in the form it was scored in.
A pipeline built by the University of Bath, published through Frontiers in AI, shows this kind of layered scoring working in a live, high-stakes setting. The system assessed credibility of tobacco-related misinformation using multiple agents working together, combined with retrieval-augmented generation and checks against outside expert judgment. It's a proof of concept, not a finished methodology, but it shows multi-agent credibility scoring holding up in a domain, public health, where getting it wrong carries real cost.
How the retrieval layer constrains what credibility scoring can see
A credibility score is only as good as the material it was run against, and if that material arrived damaged or incomplete, the score is measuring the retrieval layer's limits, not the web's actual content on the question. Stage three can only judge what stage two hands it.
A large share of pages on the web sit behind some form of bot defense, like Cloudflare challenges, browser fingerprint checks, or CAPTCHA walls. An agent that can't get past these loses access to a whole category of sources before credibility scoring ever gets a chance to look at them. And this isn't a cost-free trade-off: pages that need heavy bypass work take much longer to retrieve and cost far more per page than a simple static page, so any pipeline under pressure to move fast or spend less ends up grabbing the easy pages by default rather than the ones that actually carry the most authority.
Getting web ingestion to work at scale means moving away from scraping one URL at a time and toward pulling from sitemaps in parallel, then cleaning the raw HTML into structured Markdown, the format that keeps token costs down and preserves the structural cues an LLM needs to reason about what it's reading. A framework called ICBCBench scores source quality only for URLs that were successfully scraped, and that word "successfully" is carrying real weight. A source that's authoritative and exists but never got pulled down cleanly scores a zero, the same zero a source gets if it never existed at all. Those are very different failures, and the current scoring can't tell them apart.
One architecture built to address this problem comes from Kadoa, which generates deterministic scraper code once through an LLM, then uses AI agents to watch those scripts over time and regenerate them when a target site changes its structure. That gets reliability without having to rerun a full agent on every single extraction, and each data point it produces carries source grounding and a confidence score before it ever reaches a downstream system. Research from McGill University backs up why this kind of architecture matters: testing AI extraction across 6,000 pages spanning six domains, including Amazon, Cars.com, and Upwork, found that structured methods, meaning code generation and vision-based extraction, kept their accuracy even as page layouts changed, while simpler AI approaches produced results that swung wildly. Page structure changes quietly and often, and when it does, it causes silent retrieval failure that corrupts whatever credibility score gets built on top of it.
Retrieval infrastructure belongs inside the evaluation discipline, not off to the side as someone else's engineering problem. A reliable pipeline that returns clean, structured content is what lets credibility scoring do the job it's supposed to do. Without it, the scoring system isn't judging sources, it's judging noise.
The model-judge problem: when the scorer has the same blind spot as the agent
The natural way to automate credibility scoring is to use another LLM as the judge, but that choice brings its own serious problem into the pipeline. These judges are weakest at the exact task credibility scoring depends on most: verifying evidence.
Meta-evaluation research from 2026 found LLM judges scoring below 55% accuracy overall, with evidence verification standing out as their single weakest subtask. The scorer and the agent it's supposed to be checking share the same blind spot. It means the automated check meant to catch bad sourcing is built from the same kind of model that produces bad sourcing in the first place.
The 2026 autonomous research agents survey looked at 24 runnable systems and found that reproducibility compounds the verification gap: only a minority released the seeds or execution traces someone would need to actually rerun and confirm a result, and that same small minority were the ones reporting any novelty-verification method at all. For most systems, there's simply no artifact left behind for a judge, human or machine, to check its work against.
Benchmark integrity itself has come under direct challenge. UC Berkeley's Center for Responsible, Decentralized Intelligence showed in April 2026 that an automated scanning agent could systematically game well-known AI agent benchmarks, including SWE-bench, WebArena, OSWorld, and GAIA. If a benchmark can be gamed this reliably, the credibility scoring foundation those benchmarks are assumed to provide starts looking a lot shakier than its reputation suggests. A reliability report built from millions of tests run across thousands of production AI agents in multiple regions found an aggregate success rate only slightly above half. Automated credibility judgment is an active weak point, not a solved engineering problem sitting quietly in the background.
None of this means automated scoring should be scrapped. It means scoring needs structured verification layered on top of it rather than trusted as a finished product. NIST guidance points in this direction, recommending evaluation probes built directly into active agent workflows to run adversarial checks, with every result logged into a machine-readable audit trail. Each probe checks factual grounding, produces a clear verdict, and records the reasoning behind it for anyone who needs to review compliance later.
How web access norms are reshaping what agents can retrieve
A credibility rubric that only looks at content quality and skips over how a source was obtained is leaving out a real variable, because the rules governing how AI agents are allowed to access the web are becoming specific enough for a pipeline to actually read and act on them.
A 2025 AI Agent Index survey of deployed agents found that only a small fraction stated outright that their crawler bots follow robots.txt, while most agents offered no clear statement either way. As more agents get deployed, this gap keeps widening rather than closing.
Even among the major labs, the approach isn't consistent. OpenAI's documentation indicates that ChatGPT-User's compliance with robots.txt isn't guaranteed when a fetch is triggered directly by a user. Anthropic takes a stricter line: its documentation states that all three of its access tokens, including the user-triggered Claude-User, follow robots.txt and also honor the non-standard Crawl-delay directive.
A new standard called ai.txt is emerging to give site owners more precise control than robots.txt ever offered. A site could allow an agent to summarize an article while blocking it from pulling images, or allow use in retrieval-augmented generation while keeping the content out of training data entirely. It can also carry natural-language instructions aimed specifically at agents built to follow them. As this kind of standard spreads, a scoring system can check how a source was accessed, the same way it already checks for freshness or author accountability.
Sources
- frontiersin.org
- Source Credibility Assessment in the Realm of Information ...
- Evaluating Deep Research Agents in 2026
- Evaluating AI Agents In 2026: Benchmarks For Teams
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- SourceBench: Can AI Answers Reference Quality Web Sources?


