ERRAND: Budgeted Maintenance of Agent Memory

AuthorsBeining Wu, Zihao Ding, Jun Huang

arXiv 20262026

TL;DR

ERRAND uses a single-peaked errand index plus a deadband wage gate to price revalidation, yielding 71.9% ITT vs 61.9% for eager revalidation (+10.0pp) at the base budget.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Consolidated stores quietly decay: 79.5pp collapse in success when items flip

Deployed agents run on handed-over knowledge, but validity is fixed at write time and never budgeted for maintenance. In the main environment, success on knowledge steps drops from 85.8% while items hold to 6.3% once they flip, a 79.5pp collapse.

Frozen-policy agents using a static briefing store keep executing stale paths, flags, and price bands. The downstream consequence is high stale use and task failure, even though every item was true at handover.

HOW IT WORKS

ERRAND: Revalidation as a Priced Errand

ERRAND’s core mechanism ties drift-aware beliefs, a single-peaked errand index, a deadband gate, and a versioned lifecycle into one scheduler. These components decide which suspect items get rechecked and when.

You can think of ERRAND like a CPU scheduler with a price tag: task steps are compute, and revalidation errands are background jobs that must earn their wage. Instead of refreshing everything on a timer, ERRAND only pays to check beliefs whose answers matter enough.

This pricing lets ERRAND buy trajectory changes, not attention, using free en-route receipts and durable versions. Plain context windows or fixed-cadence refreshes cannot target decision-relevant doubt or stop spending on their own.

DIAGRAM

ERRAND Decision Flow During Deployment

This diagram shows how ERRAND routes each deployment step between working on tasks and running priced revalidation errands.

DIAGRAM

ERRAND Evaluation Pipeline Across Budgets and Worlds

This diagram shows how the experiments deploy ERRAND and baselines across different budgets, coverage gaps, models, and horizons.

PROCESS

How ERRAND Handles a Task Stream — ERRAND & Lifecycle Loop

  1. 01

    Drift-aware Belief over Atoms

    ERRAND maintains staleness belief qa for each grounded atom using closed-form relaxation under two-state drift. This belief only moves when the world provides evidence.

  2. 02

    Single-peaked Errand Index

    ERRAND computes VOIi(qi) = min(qi ℓi, (1 − qi) gi) and scales it by usage rate and durability horizon to price each suspect item.

  3. 03

    Deadband Gate with Running Wage

    ERRAND compares value per action V/c against the running wage ˆν, funding errands only when they clear the wage and otherwise holding doubt in the deadband.

  4. 04

    Versioned Lifecycle and Free Receipts

    ERRAND moves items between in-service and abeyant, writes superseding versions on repair, and uses free en-route receipts to reset beliefs without spending budget.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Closed-form Budgeted Maintenance of Memory

    ERRAND formulates belief relaxation, single-peaked value of resolving, exact deadband width |Di|, and durability horizon τi, then validates them with preregistered prediction tests where measured gains exceed forecasts by 0.6–3.6pp.

  • 02

    ERRAND Scheduler with Priced Errands

    ERRAND builds a scheduler that collects free receipts, buys the cheapest sufficient errand grade, and schedules in closed form without model calls, clearing every non-oracle policy on ITT with an average +7.1pp margin.

  • 03

    Mapping When Pricing Matters

    ERRAND shows that half-life priors collapse once drift decouples from the calendar, while ERRAND holds flat, and that the conditional premium climbs from 2.3pp to 12.9pp as coverage-gap severity increases.

RESULTS

By the Numbers

ITT success

71.9% ITT

+10.0pp over Eager revalidation at b=12

Conditional success

57.8% Cond.

+1.8pp over Eager revalidation at b=12

Spend share

11.0% steps

vs 70.7% for uncapped Eager revalidation

Average ITT margin

7.1pp ITT

ERRAND over non-oracle baselines across five settings

On the scripted tool-use worlds with drifting paths, flags, and price bands, ERRAND is evaluated against baselines under equal action budgets. The main result shows ERRAND achieving 71.9% ITT success at the base cap b=12 versus 61.9% for eager revalidation, while self-limiting spend to 11.0% of steps when uncapped.

BENCHMARK

By the Numbers

On the scripted tool-use worlds with drifting paths, flags, and price bands, ERRAND is evaluated against baselines under equal action budgets. The main result shows ERRAND achieving 71.9% ITT success at the base cap b=12 versus 61.9% for eager revalidation, while self-limiting spend to 11.0% of steps when uncapped.

BENCHMARK

Task success (%) at base budget b=12

ITT success on all errand-relevant steps at b=12 in the mid-tier 8B setting.

KEY INSIGHT

The Counterintuitive Finding

Given no cap, ERRAND stops on its own at 11.0% of steps, while uncapped eager revalidation spends 70.7% and still finishes 4.5pp behind capped ERRAND. Restraint, governed by the wage, beats aggressive refreshing even when more budget is available.

This breaks the intuition that more maintenance actions always help. ERRAND shows that pricing and ordering doubt matter more than raw spend, turning overspending into worse task success.

WHY IT MATTERS

What this unlocks for the field

ERRAND unlocks agents that can maintain long-lived consolidated memory under tight action budgets, deciding exactly which beliefs deserve rechecks and when. Builders can now deploy frozen-policy agents with briefings that stay decision-relevant, using priced errands and free receipts instead of naive periodic refreshes or unpriced self-audits.

~12 min read← Back to papers

Related papers

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: ERRAND: Budgeted Maintenance of Agent Memory

Answers use this explainer on Memory Papers.

Checking…