RoutePrism: Tracing Construction Order Effects in Agent Memory

AuthorsDong Xu, Zhangfan Yang, Jiantao Wu et al.

arXiv 20262026

TL;DR

RoutePrism uses paired construction paths and a four-condition support restoration test to show that restoring a single displaced record recovers over 60 percentage points of lost accuracy.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Construction order silently discards task evidence (restoring one record recovers over 60 points)

Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered.

Memory policies in long-horizon assistants silently drop the only supporting record for a query, causing incorrect answers while standard benchmarks report a single conflated accuracy score.

HOW IT WORKS

RoutePrism: Paired construction paths and support restoration

RoutePrism combines Task-preserving eligibility, Paired Memory Construction, Observable Diagnostics, and Support Intervention and Analysis to isolate construction-order effects under fixed policies and answer models.

Think of RoutePrism as running two different ingestion orders into the same "memory card catalog", then swapping one card back in to see if the catalog still answers correctly.

This paired-path plus four-condition design lets RoutePrism pinpoint when survivor selection or summarization loses task-relevant evidence that a plain context window view of accuracy cannot expose.

DIAGRAM

RoutePrism’s four-stage diagnostic pipeline

This diagram shows how RoutePrism runs eligibility checks, builds paired memories, computes diagnostics, and performs the four-condition support restoration intervention.

DIAGRAM

Evaluation setup across memory policies and answer models

This diagram shows how RoutePrism evaluates different memory policies on PersonaMem-32K and LongMemEval-S with multiple answer models.

PROCESS

How RoutePrism Handles a Question Route — four diagnostic stages

  1. 01

    Stage 1 Task-preserving eligibility

    RoutePrism screens each question using Task-preserving eligibility to ensure permutations of the record pool preserve the same correct answer before any construction.

  2. 02

    Stage 2 Paired Memory Construction

    RoutePrism runs Paired Memory Construction, feeding the same source pool in forward and reversed orders to the memory policy while recording exposed sources and compiled contexts.

  3. 03

    Stage 3 Observable Diagnostics

    RoutePrism computes Observable Diagnostics, including J source overlap, H answer disagreement, and Δ correctness change, plus residual route effects using repeated answer draws.

  4. 04

    Stage 4 Support Intervention and Analysis

    RoutePrism performs Support Intervention and Analysis with four conditions S+, S−, Res., and PL to test whether restoring a displaced supporting record recovers task accuracy.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Diagnostic protocol for construction-order effects

    RoutePrism introduces a paired construction protocol with Task-preserving eligibility and Paired Memory Construction, showing Dres > 0.4 residual route effects beyond answer-model sampling noise.

  • 02

    Evidence that displaced records carry task value

    RoutePrism’s Support Intervention and Analysis shows that restoring a single displaced record recovers over 60 percentage points of lost accuracy, while equal-length non-support replacements do not.

  • 03

    Distinct failure signatures across memory policies

    RoutePrism reveals that compaction, bounded recency, MemoChat-style summarization, and A-MEM each exhibit different source, context, and metadata failure signatures that no single observation layer can capture.

RESULTS

By the Numbers

Residual Dres focal

0.529

+0.036 over Recent-3 Dres 0.493

Correctness shift focal

0.444

+0.416 over Recent-3 correctness shift 0.028

Support R minus A focal

0.671

Res. 0.881 vs Abs. 0.210 on PersonaMem primary

Support R minus P Recent-3

0.750

Res. 0.917 vs Repl. 0.167 on PersonaMem primary

On PersonaMem-32K fact-recall queries, RoutePrism shows that reversing construction order under the focal compactor changes exposed sources (J = 0.057) and yields Dres = 0.529 residual route disagreement. The four-condition support restoration on PersonaMem demonstrates that RoutePrism can attribute over 60 percentage points of accuracy loss directly to displaced supporting records rather than prompt length or source count.

BENCHMARK

By the Numbers

On PersonaMem-32K fact-recall queries, RoutePrism shows that reversing construction order under the focal compactor changes exposed sources (J = 0.057) and yields Dres = 0.529 residual route disagreement. The four-condition support restoration on PersonaMem demonstrates that RoutePrism can attribute over 60 percentage points of accuracy loss directly to displaced supporting records rather than prompt length or source count.

BENCHMARK

PersonaMem whole-record restoration contrasts (Primary set)

Answer accuracy in support present, absent, restored, and replacement conditions under focal and Recent-3 policies.

KEY INSIGHT

The Counterintuitive Finding

RoutePrism shows that restoring a single displaced record can recover over 60 percentage points of lost accuracy under both compaction and bounded recency.

This is surprising because many memory systems assume that as long as some related records remain, construction order and survivor selection should not drastically change task performance.

WHY IT MATTERS

What this unlocks for the field

RoutePrism gives builders a way to audit memory policies at the evidence level, separating source selection, context compilation, and answer sampling effects.

With RoutePrism, developers can design and regression-test survivor rules, summarizers, and metadata updates to avoid silent evidence loss that was previously invisible behind a single accuracy number.

~12 min read← Back to papers

Related papers

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: RoutePrism: Tracing Construction Order Effects in Agent Memory

Answers use this explainer on Memory Papers.

Checking…