Lifestyle

ProvenanceGuard: Multiverse Computing's method for checking whether AI agents credit facts to the right source

On 29 September 2026, the Multiverse Computing team presented ProvenanceGuard. The team says it checks, after an AI agent answers, whether each claim really comes from the source the answer names. A true fact credited to the wrong source can mislead in areas such as healthcare, customer support and finance. All results are the team's own and have not been independently verified.

About 8 min read

ProvenanceGuard: Multiverse Computing's method for checking whether AI agents credit facts to the right source
Image: Mokaair (Original editorial artwork)

What happened

The Multiverse Computing team (authors listed as Antonio Tiene, Ander Alvarez Sanz and Oliver Wirjadi) published a post on the Hugging Face blog on 29 September 2026 introducing the paper "ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents". The team says the method is designed for AI agents that use multiple tools through the Model Context Protocol (MCP). MCP is a way for an AI agent to connect to and call outside tools.

As the team describes it, MCP lets an agent call search tools, view structured patient or account records, query databases and fetch metadata, then combine all of this into one answer. The problem is that common checking methods such as RAGAS faithfulness, MiniCheck, AlignScore and SummaC usually pool all the evidence before judging whether a claim is supported. They generally do not indicate which tool output actually supports that claim.

ProvenanceGuard: Multiverse Computing's method for checking whether AI agents credit facts to the right source
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

What is "cross-source conflation"?

The team calls the problem it tackles "cross-source conflation": a claim that is true somewhere in the evidence but is attributed to the wrong source. In the team's example, a customer-service agent answers "According to the account record, this plan includes a 30-day refund window." The refund window itself may be real, but it is actually stated in the policy document, not the account record. With the evidence pooled, the sentence seems grounded; looked at source by source, the attribution is wrong.

The team also gives a clinical-agent example. If a personal medication detail taken from a patient-record tool is presented in the answer as a finding from the medical literature, it becomes misleading. The team argues that in data-sensitive settings, a wrong attribution can be as harmful as a wrong fact.

How ProvenanceGuard works

According to the team, ProvenanceGuard is a "post-generation verification layer" that sits on top of a black-box MCP agent, meaning an existing agent whose inner workings it does not change. It runs after the agent produces its answer. It reads the captured MCP trace, which is the agent's record of its tool outputs and their source IDs. It requires no retraining of the agent, and it keeps track of which source each piece of evidence came from rather than merging everything into a single anonymous block. The team says it performs the following five steps in order:

  1. Split the answer into specific claims.
  2. Find the most relevant source for each claim.
  3. Check whether that source really supports the claim.
  4. Compare that source with the source the answer claims or implies.
  5. Output a source verdict for each claim, plus a pass-or-block decision for the answer as a whole.

The team says its experiments used local models. MiniLM helped find relevant sources, a DeBERTa NLI verification model checked whether a source supports a claim, and a local language model helped split answers into claims. The verifier strictly checks literal values such as numbers, dates or identifiers, so a value absent from the source will not pass just because the sentence sounds plausible. A blocked answer can go through a RARR-style repair step, which tries to rewrite the answer using the sources or to substitute safe fallback text; the result is then verified again. The team stresses that these models are only the evaluation setup, not requirements, and that switching to cloud models would need retesting and recalibration.

Test results published by the team

According to the team, the test subject was a medical agent using tools such as patient records and research articles, with 281 real traces collected. In the main test, human experts reviewed 361 claims from 40 answers set aside from the data used to develop the system. Of the 139 claims experts said should not pass, ProvenanceGuard blocked 138 and let 1 through. It also sent 67 claims that experts considered supported for review or repair. For claims whose source could be identified, the team says about 86% were matched to the correct source.

F1 here measures how well a checker blocks claims that should be blocked while avoiding unnecessary blocks. Figures published by the Multiverse Computing team, not independently verified.
VerifierReject/block F1Outputs a source ID per claim
ProvenanceGuard0.802Yes
MiniCheck0.783No
RAGAS Faithfulness0.758No
AlignScore0.662No
SummaC-ZS0.436No

The team also ran 50 controlled cases in which the named source was swapped but the supporting evidence was kept, and reports that ProvenanceGuard detected all 50. On repair, all 173 blocked answers in the full-trace test were handled, but 144 of them ended in fallback text rather than a substantive rewrite. In the reconstructed multi-source test traces, all 59 blocked answers were handled, with only 2 ending in fallback text. On performance, the team says each answer takes about half a second in its local setup.

Limitations and areas for improvement

The team itself points out weaknesses. In a harder test with several similar sources, ProvenanceGuard scored 0.846 F1 on deciding which claims to block. It correctly identified the exact source for only 50.3% of claims, and the team says distinguishing similar sources remains an important area for improvement. In addition, the setting tested was deliberately cautious: it preferred sending some supported claims for a second look over letting unsupported ones through. That is why 67 supported claims were sent for review or repair.

What it means for everyday readers

More and more AI assistants consult several systems at once before answering. For ordinary users, this research is a reminder that an AI can name a source for a fact that is true but did not come from that source. In settings such as healthcare, customer service or finance, this kind of misattribution can affect how people judge information.

The team says NVIDIA NVFlow has merged an optional grounding-verification stage for its financial agent. This stage checks completed answers against the SEC excerpts the agent retrieved and uses ProvenanceGuard's source-aware verification approach. The team also says ProvenanceGuard was presented as a poster at the Agentic AI Summit 2026 held at UC Berkeley. For ordinary users, the practical approach for now is to check the original sources an AI cites whenever the information matters.

FAQ

What is ProvenanceGuard?

According to the Multiverse Computing team, it is a verification layer that runs after an AI agent produces an answer. It checks, claim by claim, whether each claim is supported by a source and whether that source is the one the answer names.

How is it different from ordinary fact-checking tools?

The team says tools such as MiniCheck and RAGAS Faithfulness usually pool the evidence before judging whether a claim is supported, and do not indicate which tool output supports it. ProvenanceGuard records the source matched to each claim, so reviewers can see which source was checked and the verdict.

Does the AI agent need retraining to use it?

The team says no. It reads the captured MCP trace, including tool outputs and their source IDs, so it can be applied on top of an existing agent. The condition is that the agent keeps records of its tools and sources.

Are the test results reliable?

All current figures come from a post published by the research team itself, were tested mainly in a medical-agent setting, and have not yet been independently verified. The team also acknowledges that with several similar sources, only 50.3% of claims had their exact source correctly identified.

Will this affect how I use AI assistants day to day?

For now this is a research result; the application example the team mentions is NVIDIA NVFlow's financial agent. The practical takeaway for ordinary users is that the source an AI names may not be correct, so check important information against the original source yourself.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle