ProvenanceGuard Adds Source-Aware Factuality Checks to Model Context Protocol Agents
Multiverse Computing released ProvenanceGuard, an offline post-generation verification layer designed for Model Context Protocol agents. Instead of pooling tool outputs into an anonymous context, it tracks source identifiers claim-by-claim to eliminate cross-source conflation.

Impact: Medium
Why it matters
You can insert this post-generation gate into multi-tool agent pipelines to verify that LLM claims match the specific tool source they cite.
TL;DR
- 01Traditional RAG verifiers ignore tool attribution, allowing models to cite the wrong MCP tool as long as the fact exists in context.
- 02ProvenanceGuard preserves tool IDs claim-by-claim and catches 100% of deliberate source attribution swaps in benchmark tests.
- 03Offline local verifiers (MiniLM and DeBERTa) add only ~0.5s overhead per response without requiring proprietary cloud APIs.
Key facts
- Human Evaluation Packet
- 361 claims across 40 agent traces
- Invalid Claims Blocked
- 138 of 139 caught (99.3%)
- Attribution Swap Detection
- 50 of 50 caught (100%)
- Exact Source Identification
- 86% in main test, 50.3% across similar sources
- Local Pipeline Overhead
- ~0.5 seconds per answer
The Cross-Source Conflation Problem
When autonomous agents query multiple tools via the Model Context Protocol (MCP), traditional factuality scorers collapse all retrieved outputs into a single evidence pool. This creates a critical blind spot known as cross-source conflation: a claim may be supported by evidence somewhere in the execution trace, but attributed to the wrong tool or document. In high-stakes domains like finance or healthcare, quoting general policy as a specific account record causes catastrophic failures even if the text appears fully grounded.
Five-Step Verification Pipeline
ProvenanceGuard operates as an offline post-generation gate over black-box agent traces without model retraining:
1. Claim Extraction: Decomposes the full response into atomic propositional claims using a local language model. 2. Source Routing: Uses a lightweight MiniLM model to map each claim to the exact MCP tool execution ID. 3. Entailment Scoring: Evaluates grounding via a local DeBERTa natural language inference (NLI) verifier, performing strict literal checks on numbers, dates, and IDs. 4. Attribution Matching: Validates whether the attributed source explicitly cited by the model matches the tool origin. 5. Repair and Fallback: Feeds blocked outputs into an automated RARR-style revision loop to substitute unsupported claims with safe fallbacks.
Benchmark Results and Overhead
In human-annotated tests on 361 claims across 40 complex multi-tool traces, domain experts flagged 139 claims as invalid; ProvenanceGuard successfully blocked 138 of them. In a controlled benchmark with 50 deliberate attribution swaps, it caught all 50 violations. In standard local execution, MiniLM routing and DeBERTa NLI inference run in tens of milliseconds, adding approximately 0.5 seconds of total latency per answer.
Try it in 2 minutes
from dataclasses import dataclass
from typing import List
@dataclass
class MCPToolTrace:
source_id: str
tool_name: str
output_text: str
def verify_claim_provenance(claim: str, cited_source_id: str, traces: List[MCPToolTrace]) -> bool:
matched_trace = next((t for t in traces if t.source_id == cited_source_id), None)
if not matched_trace:
return False
# Run DeBERTa NLI check strictly against matched_trace.output_text
return nli_entailment_check(premise=matched_trace.output_text, hypothesis=claim)python
✓ When to use
- Complex multi-tool MCP agents that query databases, APIs, and documents in a single workflow.
- Regulated environments (clinical, financial, legal) where misattributed facts create compliance liabilities.
✕ When NOT to use
- Single-turn retrieval pipelines that query only one unified text index.
- Ultra-low latency streaming voice or chat interfaces where an extra 500ms gate is unacceptable.
What to do today
- Preserve tool invocation metadata and unique source IDs in your MCP client execution traces.
- Separate factual entailment verification from source attribution validation in agent evaluation pipelines.
- Add an automated RARR-style fallback loop to rewrite claims that fail tool-source verification.
Sources