Ugrás a tartalomra
Vissza a hírekhez
Hugging Face2026. szept. 29. 15:07eszköz

Kiszűri az AI-ügynökök forrástévesztéseit a ProvenanceGuard

A Multiverse Computing bemutatta a ProvenanceGuard nevű eszközt, amely ellenőrzi, hogy az AI-ügynökök valóban a megfelelő forrásból vették-e az adatokat.

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

A Multiverse Computing bemutatta a ProvenanceGuard nevű ellenőrző rendszert, amely az összetett feladatokat végző AI-ügynökök válaszait vizsgálja. A mai modern asszisztensek a Model Context Protocol (MCP) segítségével egyszerre több adatbázisból, keresőből és dokumentumból dolgoznak, így könnyen összekeverhetik, hogy melyik információ honnan származik. A rendszer kifejezetten ezt a forrás-összemosási hibát hivatott kiküszöbölni.

A hagyományos ellenőrzők csak azt nézik, hogy az állítás igaz-e az összesített adatok alapján, de azt nem, hogy az AI a megfelelő forráshoz társította-e azt. A ProvenanceGuard utólagos ellenőrző rétegként működik: mondatokra bontja a választ, megkeresi a leginkább releváns forrást, és összeveti a hivatkozással. Ha hibát talál, blokkolja a választ, vagy megpróbálja automatikusan javítani azt.

A fejlesztők orvosi adatokon tesztelték a rendszert, ahol a hibás forrásmegjelölés kritikus lehet. A szakértők által kiszűrt 139 hibás állításból a szoftver 138 forrástévesztést sikeresen azonosított. A kutatás részletei a Hugging Face oldalán olvashatók.

Az eredeti szöveg (Hugging Face)
The problem: supported somewhere is not the same as supported by the right source What ProvenanceGuard does Results Checking claims when sources look similar Repairing blocked answers Why this fits Multiverse Computing Tool-using LLM agents no longer read from a single retrieved passage. Through the Model Context Protocol (MCP), an agent can call a search tool, inspect a structured patient or account record, query a database, and pull metadata, then weave all of it into one answer. That m aakes the usual question of factuality more subtle than it looks. Most of the systems built to check LLM answers, from RAGAS faithfulness to fine-grained checkers like MiniCheck, AlignScore, and SummaC, ask whether a claim is supported by the available evidence once that evidence has been pooled together. In their usual form, they do not tell us which MCP tool output supports each claim, or whether that is the source the answer names. Our latest paper, ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (read it on Hugging Face, or on arXiv in the meantime), targets that gap. The failure mode we care about is one we call cross-source conflation: a claim that is true somewhere in the evidence, but attributed to the wrong source. A source-blind verifier may pass it, because the fact does exist in the pool. A source-aware verifier should not. Consider a customer support agent that answers, "According to the account record, this plan includes a 30-day refund window." The refund window may be perfectly real, but stated in a policy document, not in the account record the answer points to. Pool the two together and the claim looks supported. Keep them separate and the attribution is wrong, and in a data-sensitive setting a wrong attribution can be as damaging as a wrong fact. The same pattern shows up in a clinical agent, where a patient-specific medication detail taken from a patient-history tool becomes misleading the moment the answer presents it as a finding from the medical literature. A claim can be supported by one MCP source while the answer attributes it to another. Source-blind scoring sees support in the pooled evidence and passes it; ProvenanceGuard separately checks whether the supporting source matches the one the answer states or implies. Source: paper Figure 1. This is why faithfulness scores, useful as they are, are not enough for MCP agents. An answer carries provenance, sometimes explicitly ("according to the account record") and sometimes implicitly. ProvenanceGuard keeps that connection between claim and source available for inspection. ProvenanceGuard is a post-generation verification layer that sits on top of a black-box MCP agent. It runs after an agent produces an answer, and never collapses the evidence into one anonymous context. Instead it carries the source identity all the way through the pipeline. It reads the captured MCP trace, including the tool outputs and their source IDs, without retraining the agent. Then it does five things in sequence: it breaks the answer into specific claims, finds the source most relevant to each one, checks whether that source actually supports it, compares the source with the one the answer names or implies, and finally emits both a per-claim source verdict and a global, answer-level allow or block decision. The verification flow. Source identity is preserved through decomposition, routing, support scoring, attribution checking, and repair, rather than being pooled. Blocked answers can go through RARR-style repair and be re-verified. Source: paper Figure 2. A few of the design choices are worth calling out. For the experiments in our paper, we used local models so the captured traces could be processed in a controlled, offline setup: MiniLM helps find the relevant source, a DeBERTa NLI verifier model checks whether that source supports the claim, and a local language model helps break answers into claims. The verifier also checks literal values closely: a