Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

AuthorsEvan Chen, Shiqiang Wang, Christopher G. Brinton

arXiv 20262026

TL;DR

PLANFENCE uses dependency-scoped lineage validation at the action boundary to prevent stale-plan execution, keeping invalid actions at 0/330 while freshness-only baselines issue 330/330 invalid actions.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Stale-plan execution despite fresh state (330/330 invalid actions)

Distributed LLM-agent teams can read the latest shared facts yet still act on an obsolete plan, causing 330/330 invalid actions under freshness-only policies.

In reservation, fulfillment, and deployment workflows, an executor may hold a revised requirement but still call tools using arguments from a stale plan, leading to cancelled orders or obsolete deployments.

HOW IT WORKS

PLANFENCE: Dependency-scoped validation at the action boundary

PLANFENCE combines Exact derivation, Declared scope, Authoritative currency, and a Dependency-scoped action gate so plans cite precise public record IDs and executors validate only action-relevant keys.

You can think of PLANFENCE like a transaction commit check for tools: instead of refreshing the whole database, it validates just the specific records that authorize a pending external call.

This lineage-focused gate lets PLANFENCE guarantee that a protected action descends from current public evidence, something a plain context window or freshness-only memory cannot establish.

DIAGRAM

PLANFENCE Action-time Validation Flow

This diagram shows how PLANFENCE validates declared dependencies with owners immediately before executing a protected action, including replan and block paths.

DIAGRAM

Evaluation Pipeline and Policy Comparison

This diagram shows how PLANFENCE and baseline memory policies are evaluated across controlled workflows, update schedules, and network traces.

PROCESS

How PLANFENCE Handles a Protected Action

  1. 01

    Recording action-relevant lineage

    PLANFENCE attaches exact parent IDs to each derived public record, so requirements and role decisions record the precise versions that produced them.

  2. 02

    Declaring dependency scope

    Tool wrappers in PLANFENCE declare D(a), the semantic keys whose values can affect the protected action, forming a trusted dependency contract.

  3. 03

    Dependency-scoped action gate

    At action time, PLANFENCE traverses exact parents, queries owners for H(x), compares the dependency frontier Fa with heads, and decides authorize, replan, or block.

  4. 04

    Plan repair and revalidation

    If PLANFENCE detects a mismatch, it fetches updated records, triggers one replan, and then revalidates; any second change or incomplete response causes the action to fail closed.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Identifying stale-plan execution and lineage validity

    PLANFENCE formalizes stale-plan execution and lineage validity, showing that freshness-only checks issue 330/330 invalid actions while exact-lineage policies issue 0/330 invalid actions.

  • 02

    Design of dependency-scoped PLANFENCE gate

    PLANFENCE introduces a dependency-scoped action gate that binds plans to exact public inputs, validates only tool-declared dependencies, and permits one replan before failing closed.

  • 03

    Coordination stall boundary across policies

    PLANFENCE maps when proactive synchronization or action-time validation is cheaper, showing lower stall and traffic than batched all-key validation across 8–128 key settings.

RESULTS

By the Numbers

Invalid actions per issued

0/330 actions

-330/330 versus Local replica

Availability per scheduled

330/330 actions

matches Centralized lineage availability

Coordination stall

230.8 ms/action

177.8 ms lower than Centralized lineage on AT&T

Distributed traffic

8.1 KiB/action

7.5 KiB lower than Batched all-key validation at 8 keys

These numbers come from the eight-key compact-state workload on the pinned AT&T trace at high churn, with 42 updates and 11 protected actions per episode. They show that PLANFENCE can maintain zero invalid actions while reducing coordination stall and traffic compared with freshness-only baselines and safe all-key validation policies.

BENCHMARK

By the Numbers

These numbers come from the eight-key compact-state workload on the pinned AT&T trace at high churn, with 42 updates and 11 protected actions per episode. They show that PLANFENCE can maintain zero invalid actions while reducing coordination stall and traffic compared with freshness-only baselines and safe all-key validation policies.

BENCHMARK

Safety and coordination cost comparison on AT&T (eight-key compact state, high-update workload)

Invalid actions and coordination stall per action for PLANFENCE and baseline memory policies.

KEY INSIGHT

The Counterintuitive Finding

Even with perfectly fresh replicated state, Local replica and Owner-head freshness each issue 330/330 actions from obsolete plans, while PLANFENCE issues 0/330 invalid actions.

This is surprising because many systems assume that fresh memory is enough; PLANFENCE shows that without explicit derivation lineage, executors can still act on stale authorization despite current state.

WHY IT MATTERS

What this unlocks for the field

PLANFENCE unlocks lineage-safe tool execution in distributed LLM-agent teams, ensuring that protected actions are authorized by current, action-relevant public records.

Builders can now design agent workflows that tolerate high churn and large shared keyspaces without global synchronization, validating only the dependencies that matter for each external call.

~12 min read← Back to papers

Related papers

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Agent Memory

ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents

Xiaohui Zhang, Zequn Sun et al.

· 2026

ActMem transforms dialogue history into atomic facts via Memory Fact Extraction, groups them with Fact Clustering, links them through a Memory KG Construction module, and uses Counterfactual-based Retrieval and Reasoning for action-aware answers. On ActMemEval, ActMem reaches 76.52% QA accuracy with DeepSeek-V3, beating LightMem’s 63.97% by 12.55 points and NaiveRAG’s 61.54%.

Questions about this paper?

Paper: Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

Answers use this explainer on Memory Papers.

Checking…