Skip to main content
Back to Blog
AI/MLSecurityEnterprise
26 September 20268 min readUpdated 29 September 2026

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

Getting the Source Right, Not Just the Fact: Source Aware Verification for MCP Agents Tool using LLM agents no longer rely on a single retrieved passage. Through the , an agent...

By AI Engineering Team

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

Tool-using LLM agents no longer rely on a single retrieved passage. Through the Model Context Protocol (MCP), an agent can call a search tool, inspect a structured patient or account record, query a database, and retrieve metadata before combining the results into one answer.

This makes factuality more nuanced. Many systems used to evaluate LLM answers, including RAGAS faithfulness, MiniCheck, AlignScore, and SummaC, determine whether a claim is supported by the available evidence after that evidence has been pooled. In their standard forms, they generally do not identify which MCP tool output supports each claim, or verify that the supporting output matches the source named in the answer.

The paper ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents, available on arXiv, addresses this gap. It focuses on cross-source conflation: a claim is true somewhere in the evidence, but the answer attributes it to the wrong source. A source-blind verifier may accept the claim because the fact exists in the evidence pool. A source-aware verifier should reject the attribution.

Supported Somewhere Is Not the Same as Supported by the Right Source

Consider a customer-support agent that says, “According to the account record, this plan includes a 30-day refund window.” The refund period may be valid, but the information could come from a policy document rather than the account record referenced by the answer.

When both documents are pooled, the claim appears supported. When the sources remain separate, the attribution is incorrect. In a data-sensitive setting, incorrect attribution can be as damaging as an incorrect fact. A similar issue can arise in a clinical agent if a patient-specific medication detail from a patient-history tool is presented as a finding from medical literature.

A claim can be supported by one MCP source while the answer attributes it to another. Source-blind scoring sees support in the pooled evidence and accepts it. ProvenanceGuard separately checks whether the supporting source matches the source stated or implied by the answer. Source: paper Figure 1.

Faithfulness scores remain useful, but they do not fully capture this problem in MCP agents. Answers can include provenance explicitly, such as “according to the account record,” or implicitly through their wording. ProvenanceGuard preserves the connection between each claim and its source so that it can be inspected.

What ProvenanceGuard Does

ProvenanceGuard is a post-generation verification layer designed to operate on top of a black-box MCP agent. It runs after the agent generates an answer and does not merge the evidence into one anonymous context. Instead, it carries source identity through the verification pipeline.

The system reads the captured MCP trace, including tool outputs and source IDs, without retraining the agent. It then performs five steps:

  1. Breaks the answer into specific claims.
  2. Finds the source most relevant to each claim.
  3. Checks whether that source supports the claim.
  4. Compares the supporting source with the source named or implied by the answer.
  5. Produces both a per-claim source verdict and a global allow-or-block decision for the answer.

The verification flow preserves source identity through decomposition, routing, support scoring, attribution checking, and repair. Blocked answers can go through RARR-style repair before being verified again. Source: paper Figure 2.

For the reported experiments, the system used local models so that captured traces could be processed in a controlled, offline environment. MiniLM helped identify the relevant source, while a DeBERTa NLI verifier model checked whether that source supported the claim. A local language model decomposed answers into claims.

The verifier also checks literal values. A number, date, or identifier absent from the source cannot pass solely because the sentence appears plausible. A calibration step combines these signals. If an answer is blocked, a RARR-style repair step can attempt a source-grounded revision or provide a safe fallback. The revised answer is then verified again.

These models are the configuration evaluated in the paper, not requirements of ProvenanceGuard. The same claim, source, and decision stages could be adapted to hosted models, although a new setup would require separate testing and calibration. The reported results come from the local configuration. Its conservative policy is intended for data-sensitive review, where checking the source can be more important than returning an answer as quickly as possible.

Results

The evaluation used answers from a medical agent that had accessed patient records, research articles, and other tools. The dataset contained 281 real traces. Medicine provides a useful test case because information from a patient's record and information from general research represent different sources.

For the main evaluation, human experts reviewed 361 claims from 40 answers that had been held out from the data used to develop the system.

Experts determined that 139 claims should not pass. ProvenanceGuard caught 138 of them and allowed one through. It also held 67 claims that experts considered supported, sending them for review or repair. This result reflects the conservative configuration: it accepts additional review for some supported claims in exchange for reducing the number of unsupported claims that pass.

For claims with an identifiable source, the system selected the correct source approximately 86% of the time in this evaluation.

Four other support checkers were tested on the same claims. ProvenanceGuard achieved the highest score on the paper's measure of detecting claims that should be blocked while avoiding unnecessary blocks. The other systems did not identify which tool output supported each claim. ProvenanceGuard records that relationship, allowing reviewers to inspect the source checked for each claim and the resulting decision.

VerifierReject/block F1Emits claim-to-source ID
ProvenanceGuard0.802Yes
MiniCheck0.783No
RAGAS Faithfulness0.758No
AlignScore0.662No
SummaC-ZS0.436No

Binary support metrics on the same held-out claim packet. ProvenanceGuard matched or exceeded the source-blind baselines on blocking while also producing per-claim source verdicts. Source: paper abstract and Table III.

Checking Claims When Sources Look Similar

A separate, more difficult evaluation used several similar sources. ProvenanceGuard achieved an F1 score of 0.846 for deciding which claims to block, but identified the exact source correctly for 50.3% of claims. Distinguishing among similar sources remains a challenge.

The researchers also conducted a controlled wrong-attribution test. In 50 cases, they changed the source named by the answer while leaving the supporting evidence unchanged. ProvenanceGuard detected all 50 source swaps. This demonstrates that it can identify clear source errors, while the multi-source evaluation highlights the difficulty of selecting the correct source when several are plausible.

Repairing Blocked Answers

Blocking is most useful when a blocked answer can be repaired or replaced. Connected to the RARR-style repair loop, the full-trace run resolved all 173 blocked answers. Of those, 144 ended in fallback text rather than a substantive rewrite, meaning the system avoided producing an unverifiable answer.

On reconstructed multi-source test traces, a new repair run resolved all 59 initially blocked answers, with only two ending in terminal fallbacks.

In the reported local configuration, the offline gate added roughly half a second per answer. The NLI and routing calls themselves took tens of milliseconds.

Application to Multitool Agents

As agents move from single-passage retrieval-augmented generation to multi-tool MCP systems, identifying the source of a fact becomes part of factuality rather than a secondary detail. ProvenanceGuard exposes that relationship claim by claim.

The medical study is one use case. The same approach can be adapted to other domains when an agent's trace preserves its tools and source IDs.

One example is NVIDIA NVFlow, which merged an optional grounding-verification stage for its finance agent. The stage checks completed answers against SEC excerpts retrieved by the agent and saves separate decisions without changing the original rollout or training data. The NVFlow contribution uses ProvenanceGuard's source-aware verification approach. The repair loop described above belongs to the broader research system.

ProvenanceGuard was also presented as a poster at the Agentic AI Summit 2026 at UC Berkeley.