A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

AuthorsXiaoyang Li, Yiqi Wang, Chencheng Zhu et al.

arXiv 20262026

TL;DR

Correlated Promotion Benchmark (CPB) shows that gating on declared source type is the only tested admission policy that keeps false adoption as low as 0.06–0.09 while still answering, whereas independence-style deduplication rejects more true than false claims.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Shared memory admits copied false claims and reinforces them (false adoption up to 0.97–0.99)

Shared agent memory misreads repeated claims as independent evidence, so copied or paraphrased false beliefs are admitted and reinforced as if they were corroborated.

Correlated Promotion Benchmark (CPB) shows that once an uncontested false belief enters shared memory, the consumer asserts it in 0.97–0.99 of probes, spreading false adoption across the whole team.

HOW IT WORKS

Correlated Promotion Benchmark — CPB-Static and CPB-Live as an instrumented admission pipeline

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to measure admission decisions and their downstream effects.

You can think of CPB like a lab where shared memory is a central database, agents are lab technicians, and the governance rule is a strict lab supervisor deciding which observations become official records.

By fixing lineage in scenarios and logging every write and retrieval, CPB enables analyses of admission policies that a plain context window cannot, including reach, exposure, and adoption of false beliefs over time.

DIAGRAM

CPB-Live interaction loop — retrieve, emit, decide, and probe adoption

This diagram shows how CPB-Live runs multi-agent teams over a shared store, logs retrievals and writes, and then probes a consumer that answers only from shared memory.

DIAGRAM

CPB evaluation pipeline — CPB-Static vs CPB-Live

This diagram shows how CPB-Static assembles a frozen benchmark with a gold admission rule, and how CPB-Live executes scenarios with logged lineage and consumer probes.

PROCESS

How Correlated Promotion Benchmark Handles a CPB-Live Episode

  1. 01

    Scenario families and feeds

    Correlated Promotion Benchmark (CPB) instantiates scenario families like Correlated agreement, Concurrent conflict, Mixed scenarios, and Dual source with authored feeds and fixed lineage.

  2. 02

    Team over shared store

    CPB runs six-agent teams for six rounds over a shared store, where agents retrieve current beliefs and emit candidate claims each round.

  3. 03

    Admission policy actions

    CPB applies admission policies such as Governance, Independence vote, and Majority vote to map each claim to PROMOTE, REQUEST EVIDENCE, KEEP PRIVATE, or ABSTAIN.

  4. 04

    Consumer probe and grading

    At round six, CPB queries a Consumer that answers solely from the store, then grades reach, exposure, and adoption of false beliefs using the scenario’s truth and lineage.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    An instrument: CPB

    Correlated Promotion Benchmark (CPB) introduces CPB-Static and CPB-Live with adapters, a gold admission rule, and lineage collapse over public annotated sources and authored multi-agent envelopes.

  • 02

    Diagnosis of admission mechanisms

    CPB shows that reimplemented voting, collapse, and judge mechanisms, plus components from Mem0 and A-MemGuard, still admit false claims under rewording or authoritative source typing.

  • 03

    Measurements of the store itself

    CPB measures how often a Consumer adopts false shared beliefs, how competing truths change adoption, and how rewording and arrival order affect admission across agent families.

RESULTS

By the Numbers

Damage share Qwen

0.152

-0.848 vs Share all

False adoption

0.06–0.09

vs 0.22–0.47 for other answering policies

False promotion Static

0.077

-0.151 to -0.208 vs 0.228–0.285 policies

Uncontested adoption

0.97–0.99

Consumer asserts uncontested false beliefs almost always

Correlated Promotion Benchmark (CPB) evaluates admission policies on CPB-Live and CPB-Static, measuring damage share, false admissions, and adoption. The main result shows that the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while other answering policies reach 0.22–0.47 false adoption and majority vote often matches share-all’s damage.

BENCHMARK

By the Numbers

Correlated Promotion Benchmark (CPB) evaluates admission policies on CPB-Live and CPB-Static, measuring damage share, false admissions, and adoption. The main result shows that the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while other answering policies reach 0.22–0.47 false adoption and majority vote often matches share-all’s damage.

BENCHMARK

CPB-Live damage share by policy on Qwen3.5-9B

Damage share (Dmg) on CPB-Live relative to Share all for Qwen3.5-9B.

KEY INSIGHT

The Counterintuitive Finding

Correlated Promotion Benchmark (CPB) finds that independence-style policies reject 0.051–0.127 of true claims while still admitting 0.119–0.306 of false ones, giving damage shares around 0.243–0.289.

This is counterintuitive because deduplication is expected to protect against copied falsehoods, yet CPB shows it disproportionately suppresses true claims while leaving many false beliefs admitted and retrievable.

WHY IT MATTERS

What this unlocks for the field

Correlated Promotion Benchmark (CPB) gives researchers a controlled way to study epistemic admission with fixed lineage, logged retrievals, and graded adoption across realistic multi-agent scenarios.

Builders can now systematically test governance rules, deduplication schemes, and judge-based filters for shared agent memory, observing how false beliefs propagate and how source-type gating or lineage-aware policies change downstream behavior.

~12 min read← Back to papers

Related papers

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Agent Memory

ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents

Xiaohui Zhang, Zequn Sun et al.

· 2026

ActMem transforms dialogue history into atomic facts via Memory Fact Extraction, groups them with Fact Clustering, links them through a Memory KG Construction module, and uses Counterfactual-based Retrieval and Reasoning for action-aware answers. On ActMemEval, ActMem reaches 76.52% QA accuracy with DeepSeek-V3, beating LightMem’s 63.97% by 12.55 points and NaiveRAG’s 61.54%.

Questions about this paper?

Paper: A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Answers use this explainer on Memory Papers.

Checking…