Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems

AuthorsYi Ting Shen, Kentaroh Toyoda, Alex Leung

arXiv 20262026

TL;DR

Revocation Enforcement Study shows that a simple retrieval-time validity guard can cut unsafe actions from 43.1% to 0% on exposed memory systems.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Revoked policies still drive actions in 43.1% of trials

Revocation Enforcement Study finds that on two major memory systems, revoked policies are still retrieved and used, causing unsafe actions in 43.1% of trials.

These failures occur in long-term agent memory for organizational policies, where revoked rules like PII export or backup deletion still reach agents and override safety constraints.

HOW IT WORKS

Revocation Enforcement Study — measuring soft revocation and a retrieval guard

Revocation Enforcement Study centers on a two-phase setup phase, experimental phase, and a retrieval-time guard that inspects records from Graphiti, mem0, langmem, cognee, and Zep.

You can think of the guard like a firewall between RAM and disk: it lets all reads pass through but strips out stale or conflicting entries before they reach the agent.

This retrieval-time validity check lets Revocation Enforcement Study enforce current policy versions in memory, something a plain context window and similarity-only retrieval cannot guarantee.

DIAGRAM

Single-read interaction between user, memory store, guard, and agent

This diagram shows how Revocation Enforcement Study runs each trial: a user query triggers retrieval, the guard filters revoked or conflicting records, and the agent chooses an action.

DIAGRAM

Evaluation pipeline across scenarios, models, and defense conditions

This diagram shows how Revocation Enforcement Study evaluates exposure and unsafe-action rates over nine scenarios, nine models, and six defense conditions.

PROCESS

How Revocation Enforcement Study Handles a Single-read Scenario

  1. 01

    Setup phase

    Revocation Enforcement Study uses the setup phase to insert a revoked policy v1 and its replacement v2 into Graphiti, mem0, langmem, cognee, or Zep under direct or indirect insertion.

  2. 02

    Experimental phase

    In the experimental phase, Revocation Enforcement Study issues an ordinary query qs, runs default retrieval R(qs,M,k), and passes txt(r) to the agent without status metadata.

  3. 03

    Defense conditions

    Revocation Enforcement Study applies store-level filter, prompt hardening, output filter, filter plus prompt, or the guard to see how each defense changes Pr[a = a−].

  4. 04

    Outcome measures

    Revocation Enforcement Study computes the exposure rate Pr[ER = 1] and the unsafe-action rate Pr[a = a−], decomposed as Pr[ER = 1] · Pr[a = a− | ER = 1].

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Measurement of revocation enforcement in five memory systems

    Revocation Enforcement Study systematically measures Graphiti, mem0, langmem, cognee, and Zep, showing that revoked policies are still returned and ranked first on two systems, with unsafe actions in up to 44.2% of trials.

  • 02

    Characterization of three revocation failure modes

    Revocation Enforcement Study identifies failures where revocation is never recorded, recorded but invisible, or recorded and visible but not enforced at retrieval, and shows Zep exposes no status to callers.

  • 03

    Design and evaluation of a retrieval-time guard

    Revocation Enforcement Study proposes a guard that intercepts reads, checks revocation marks or conflicts, and matches the store-level filter with 0/1,620 unsafe actions where status is exposed, while still reducing unsafe actions when no mark exists.

RESULTS

By the Numbers

Unsafe-action rate Graphiti

44.2%

+2.1 percentage points over mem0 (exp.)

Unsafe-action rate mem0 (exp.)

42.1%

699/1,620 unsafe trials pooled across exposed systems

Exposure rate Graphiti

100.0%

81/81 scenarios return revoked policy ranked first

Unsafe-action rate with rule

12.9%

drops to 4.6% when excluding backup deletion outlier

Revocation Enforcement Study evaluates nine scenarios and nine models across five memory systems, measuring how often revoked policies are retrieved and cause unsafe actions. The main result shows that default retrieval on Graphiti and mem0 (expiry override) returns revoked policies in every scenario and leads to unsafe actions in 42.1–44.2% of trials, while retrieval filtering or the guard eliminates these failures.

BENCHMARK

By the Numbers

Revocation Enforcement Study evaluates nine scenarios and nine models across five memory systems, measuring how often revoked policies are retrieved and cause unsafe actions. The main result shows that default retrieval on Graphiti and mem0 (expiry override) returns revoked policies in every scenario and leads to unsafe actions in 42.1–44.2% of trials, while retrieval filtering or the guard eliminates these failures.

BENCHMARK

Unsafe-action rates under no defense across systems

Unsafe-action rate Pr[a = a−] under no defense for exposed versus non-exposed configurations.

KEY INSIGHT

The Counterintuitive Finding

Revocation Enforcement Study shows that even with explicit prompt rules, agents still take unsafe actions in 12.9% of trials, reaching 46.1% in backup deletion.

This is counterintuitive because many builders assume a strong system prompt can override stale memory, yet revoked policies still dominate decisions when retrieval surfaces them.

WHY IT MATTERS

What this unlocks for the field

Revocation Enforcement Study unlocks a practical retrieval-time validity guard that can sit between any agent and memory backend without vendor changes.

With this, builders can deploy long-term agents whose actions respect current policies, avoiding revoked rules silently reappearing through memory retrieval.

~12 min read← Back to papers

Related papers

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

Questions about this paper?

Paper: Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems

Answers use this explainer on Memory Papers.

Checking…