MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents

AuthorsJiajun Dong, Yutao Hu, Fengrui Fan et al.

arXiv 20262026

TL;DR

MemArbiter uses function-aware memory banks plus a temporal presentation gate to close the Memory-Action Gap, reaching 92.54% SR@50 on ALFWorld (+25.38pp over Flat Retrieval).

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Accessible Memories Still Fail to Guide Actions: The Memory Action Gap

MemArbiter targets the Memory Action Gap, where action relevant information is accessible yet “fails to guide action selection” despite being retained.

In long horizon ALFWorld tasks, flat memory organization mixes heterogeneous Goal, Task State, and Episodic items, causing repeated failures and low success rates even with large memory budgets.

HOW IT WORKS

MemArbiter: Function Aware Memory Arbitration

MemArbiter combines Memory Banks, a Candidate Writer, Dual Band Memory Representation, Decision Relevance Signals, a Temporal Presentation Gate, and Prompt Assembly to control decision time memory salience.

You can think of MemArbiter like a smart RAM controller plus a card catalog, deciding which “cards” stay fully visible and which become compact cues.

This KEY_MECHANISM of stateful, function aware presentation lets MemArbiter influence actions beyond a plain context window that only stores or retrieves text blocks.

DIAGRAM

Decision Time Flow Through MemArbiter

This diagram shows how MemArbiter processes each step in the ReAct loop from observation to action using its arbitration pipeline.

DIAGRAM

ALFWorld Evaluation and Ablation Design

This diagram shows how MemArbiter is evaluated against Flat Recency and Flat Retrieval on ALFWorld, including ablation variants.

PROCESS

How MemArbiter Handles a ReAct Decision Step

  1. 01

    State Parsing

    MemArbiter parses the task goal, current observation, previous action, and outcome to form the decision context c_t and identify subgoals and information needs.

  2. 02

    Memory Banks and Candidate Writes

    Using the Candidate Writer, MemArbiter decomposes new evidence into atomic items and updates the Goal, Task State, Constraint, Episodic, and Reference Memory Banks.

  3. 03

    Decision Relevance Signals and Temporal Presentation Gate

    MemArbiter computes bank level demand and item level relevance, then the Temporal Presentation Gate assigns each item to focal, ambient, or hidden states using relevance history.

  4. 04

    Prompt Assembly

    MemArbiter groups focal and ambient items by Memory Bank, orders them by relevance, and assembles the structured memory prompt P_t for the next action generation.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Memory Action Gap Definition

    MemArbiter formalizes the Memory Action Gap as a post access failure where accessible information “fails to guide action selection,” distinct from forgetting or retrieval failure.

  • 02

    Function Aware Memory Arbitration

    MemArbiter introduces five functional Memory Banks plus Dual Band Memory Representation and a Temporal Presentation Gate to dynamically control timing and salience under budgets.

  • 03

    Improved Long Horizon Agent Performance

    On ALFWorld, MemArbiter reaches 82.84% and 92.54% SR@50 under 500 and 750 token budgets, improving over Flat Retrieval by 20.90 and 25.38 percentage points.

RESULTS

By the Numbers

SR@50 500 tokens

82.84%

+20.90 over Flat Retrieval

SR@50 750 tokens

92.54%

+25.38 over Flat Retrieval

Pick Two 500 tokens

58.82%

Flat Recency and Flat Retrieval both 0.00%

One step recovery 750 tokens

65.3%

+34.8 over Flat Retrieval

On the 134 unseen ALFWorld tasks, which require multi step navigation and manipulation, these numbers show MemArbiter turning accessible memories into effective actions. The MAIN_RESULT demonstrates that function aware arbitration is crucial when memory budgets are fixed but trajectories are long.

BENCHMARK

By the Numbers

On the 134 unseen ALFWorld tasks, which require multi step navigation and manipulation, these numbers show MemArbiter turning accessible memories into effective actions. The MAIN_RESULT demonstrates that function aware arbitration is crucial when memory budgets are fixed but trajectories are long.

BENCHMARK

Success rate by task type on the 134 task ALFWorld unseen split

SR@50 on ALFWorld unseen tasks under a 500 token memory prompt budget.

BENCHMARK

Ablation results on ALFWorld under the 500 token memory budget

SR@50 for MemArbiter and its ablations on ALFWorld unseen tasks.

KEY INSIGHT

The Counterintuitive Finding

Even when the memory budget increases from 500 to 750 tokens, Flat Retrieval only rises from 61.94% to 67.16% SR@50, a modest 5.22 point gain.

This is surprising because we usually expect more context to help, but MemArbiter shows that without function aware arbitration, extra tokens barely translate into better decisions.

WHY IT MATTERS

What this unlocks for the field

MemArbiter unlocks agents that treat memories as typed decision tools, not just text logs, preserving goals, constraints, and episodes with appropriate salience over time.

Builders can now design long horizon agents whose memory modules actively arbitrate which facts, states, and experiences shape each action, rather than hoping retrieval alone suffices.

~12 min read← Back to papers

Related papers

Memory Architecture

A Control Architecture for Training-Free Memory Use

Yanzhen Lu, Muchen Jiang et al.

· 2026

TAG routes low-confidence steps to uncertainty-based routing, filters them with guarded acceptance with rollback, chooses between bank selection across rule and exemplar memory, and prunes via evidence-based retirement inside a unified control loop. On SVAMP and ASDiv, TAG reaches 81.0% and 85.2% accuracy, improving over the 74.0% and 77.5% no-memory baselines while a compute-matched Retry baseline stays flat.

Questions about this paper?

Paper: MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents

Answers use this explainer on Memory Papers.

Checking…