TEPA: Revoking Stale Memories for Conflict-Robust Language Agents

AuthorsYan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang

arXiv 20262026

TL;DR

TEPA uses conflict-keyed lifecycle revocation to keep only valid memories active, reaching 0.950 success under full reversal vs 0.210 for append-only memory.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Memory Pollution: Append-only Agents Collapse Under Reversal (append-only 0.210 vs no memory 0.309)

Append-only and last-write-wins memory fall to 0.210 success during full reversal, below the no-memory baseline of 0.309, revealing memory pollution.

In this regime, language agents retrieve stale evidence as prompt context, causing conflicting facts to dominate decisions and making persistent memory actively harmful.

HOW IT WORKS

TEPA: Revocable Evidence-Memory with Lifecycle States

TEPA represents observations as keyed precedents with explicit lifecycle state, separates an active set from a revoked archive, and uses a retriever plus trial-validated promotion to manage validity.

You can think of TEPA like a card catalog where each card tracks a fact’s key, success statistics, and status, and stale cards are moved to an audit drawer instead of the main shelf.

This conflict-keyed lifecycle mechanism lets TEPA falsify and revoke outdated facts, something a plain context window or append-only cache cannot do while still preserving history for audit.

DIAGRAM

TEPA Lifecycle Update and Revocation Flow

This diagram shows how TEPA updates precedent states, increments support or conflict counts, and revokes stale precedents during each episode.

DIAGRAM

Evaluation Pipeline Across Drift and Preference Benchmarks

This diagram shows how TEPA is evaluated on controlled drift, executable drift, preference updates, and MemoryAgentBench with shared baselines.

PROCESS

How TEPA Handles an Episode Stream Interaction

  1. 01

    Episode Stream and Context

    TEPA receives an episode et with visible context ct and uses the retriever to pull relevant precedents from the active set At.

  2. 02

    Policy Execution and Outcome

    Conditioned on ct and retrieved precedents, TEPA’s policy πm selects action a(m)t and observes binary outcome y(m)t plus evidence xt.

  3. 03

    Conflict Key Extraction

    TEPA applies extractors κ and ν to xt to obtain conflict key kt and value vt, then updates keyed precedents with support or conflict counts.

  4. 04

    Lifecycle Update and Revocation

    Using the Beta-Bernoulli posterior q(p), TEPA performs lifecycle update, revoking precedents into the revoked archive Rt or promoting hypotheses to the active set.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Formulating Memory Pollution and Pollution Index

    TEPA defines memory pollution as a validity-state failure and introduces a phase-wise memory pollution index MPIm,ϕ, showing append-only success dropping to 0.210 vs 0.309 for no memory.

  • 02

    Conflict-Keyed Revocable Evidence Memory

    TEPA proposes keyed precedents with explicit lifecycle states and revocation, separating an active set At from a revoked archive Rt for falsifiable long-term memory.

  • 03

    Lifecycle Evaluation Across Four Conflict Settings

    TEPA is evaluated on controlled drift, executable drift, preference-update streams, and MemoryAgentBench SH6k, reaching 0.950 reversal success and 0.890 substring exact match on SH6k.

RESULTS

By the Numbers

Full-reversal success

0.950

+0.740 over Last-write-wins

Executable reversal success

0.950

+0.747 over Append-only

Preference full-reversal success

0.872

+0.034 over No memory

MemoryAgentBench SH6k substring EM

0.890

+0.307 over Append-only

These metrics come from controlled hidden-regime drift, real file-backed executable drift, preference-update streams, and MemoryAgentBench SH6k, which test stale-conflict consolidation. The 0.950 reversal success and 0.890 substring exact match show TEPA keeps memory useful where append-only and last-write-wins become worse than no memory.

BENCHMARK

By the Numbers

These metrics come from controlled hidden-regime drift, real file-backed executable drift, preference-update streams, and MemoryAgentBench SH6k, which test stale-conflict consolidation. The 0.950 reversal success and 0.890 substring exact match show TEPA keeps memory useful where append-only and last-write-wins become worse than no memory.

BENCHMARK

Main Evidence Summary Across Conflict Settings

Success rate across controlled drift, executable drift, preference updates, and MemoryAgentBench SH6k.

KEY INSIGHT

The Counterintuitive Finding

Under full reversal, append-only and last-write-wins memory drop to 0.210 success, below the no-memory baseline of 0.309, yielding a pollution index of 0.318.

This is surprising because persistent memory is usually assumed to help agents, yet TEPA shows that without revocation, stored experience can be strictly worse than forgetting.

WHY IT MATTERS

What this unlocks for the field

TEPA unlocks long-term memory that can falsify, audit, and later re-promote evolving knowledge by tracking validity state per conflict key.

Builders can now design agents whose memories remain trustworthy under regime shifts and preference changes, instead of relying on brittle append-only caches or global resets.

~12 min read← Back to papers

Related papers

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

Questions about this paper?

Paper: TEPA: Revoking Stale Memories for Conflict-Robust Language Agents

Answers use this explainer on Memory Papers.

Checking…