MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents

AuthorsYongxian Wei, Yilin Zhao, Runxi Cheng et al.

arXiv 20262026

TL;DR

MemAgent uses content aware routing over heterogeneous memory providers to raise average accuracy to 75.7% on GAIA WebWalkerQA and xBench DS, +10.0 points over No Memory.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

No single memory works across benchmarks 18.6 percent oracle gap

MemAgent’s study of 13 memory methods shows that no single provider dominates, while an oracle over all providers reaches 89.6% accuracy versus 71.0% for the best single provider.

This means agents with a fixed memory representation waste complementary strengths across tasks, leaving up to 18.6 percentage points of accuracy unrealized and limiting long horizon and cross task performance.

Without heterogeneous routing, agents remain largely stateless across tasks, preventing continual improvement and amortization of effort on GAIA, WebWalkerQA, and xBench DS.

HOW IT WORKS

MemAgent — content aware routing over heterogeneous memory providers

MemAgent combines a content aware routing architecture, a training data synthesis pipeline, InsightGraph, and a short term memory provider to manage BEGIN, IN, and END memory decisions.

You can think of MemAgent as a smart operating system that chooses between different memory disks and a RAM buffer, instead of relying on one oversized but clumsy drive.

This KEY_MECHANISM lets MemAgent preview, gate, and selectively store experiences so agents reuse the right memories at the right time, beyond what a plain context window can encode.

DIAGRAM

Three phase memory lifecycle routing in MemAgent

This diagram shows how MemAgent routes memory at BEGIN, IN, and END phases between the task agent and heterogeneous providers.

DIAGRAM

MemAgent training data synthesis and reward filtering

This diagram shows how MemAgent synthesizes supervision via forced exploration, per decision rewards, and reward filtered supervised fine tuning.

PROCESS

How MemAgent Handles a Task — BEGIN IN END lifecycle

  1. 01

    BEGIN phase content aware probing

    MemAgent uses the content aware routing architecture to probe all providers, including InsightGraph and other long term memory providers, before selecting one for retrieval.

  2. 02

    IN phase short term memory gating

    During each step, MemAgent consults the short term memory provider and decides whether to inject a working memory summary or skip injection.

  3. 03

    END phase multi provider storage

    After the trajectory finishes, MemAgent uses the content aware routing architecture to select a subset of providers for storing the trajectory for future reuse.

  4. 04

    Training data synthesis pipeline

    Offline, MemAgent’s training data synthesis pipeline runs forced exploration, per decision reward scoring, and reward filtered supervised fine tuning to learn routing policies.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Rethinking agent memory with heterogeneous providers

    MemAgent unifies 13 memory methods as providers, adds InsightGraph and a short term memory provider, and shows an oracle reaches 89.6% versus 71.0% for the best single provider.

  • 02

    Three phase routing formulation BEGIN IN END

    MemAgent formulates memory as routing decisions at BEGIN, IN, and END, implemented by a content aware routing architecture that previews, gates, and selectively stores experiences.

  • 03

    Training data synthesis for routing without labels

    MemAgent introduces a training data synthesis pipeline combining forced exploration, LLM as Judge rewards, and downstream reuse signals to supervise routing without ground truth labels.

RESULTS

By the Numbers

GAIA Overall

69.7%

+10.9 over No Memory

WebWalkerQA

79.4%

+10.0 over No Memory

xBench DS

78.0%

+9.0 over No Memory

Avg.

75.7%

+10.0 over No Memory

On GAIA, WebWalkerQA, and xBench DS, which test multi step tool use and deep search, MemAgent raises average accuracy from 65.7% to 75.7%, proving that heterogeneous routing yields a 10.0 percentage point gain over a stateless No Memory agent.

BENCHMARK

By the Numbers

On GAIA, WebWalkerQA, and xBench DS, which test multi step tool use and deep search, MemAgent raises average accuracy from 65.7% to 75.7%, proving that heterogeneous routing yields a 10.0 percentage point gain over a stateless No Memory agent.

BENCHMARK

Accuracy on GAIA Overall WebWalkerQA and xBench DS

Accuracy (%) of MemAgent versus No Memory, best single provider, and PromptRoute across three benchmarks.

KEY INSIGHT

The Counterintuitive Finding

MemAgent shows that an oracle over 13 providers reaches 89.6% accuracy, while the best individual provider only reaches 71.0%.

This breaks the assumption that designing a single superior memory representation is enough, revealing that routing across heterogeneous memories is crucial.

Even sophisticated single providers like AWM and ExpeL lead on only one benchmark each, contradicting expectations of a universal memory format.

WHY IT MATTERS

What this unlocks for the field

MemAgent unlocks adaptive, content aware orchestration of diverse memory systems, raising average accuracy by 10.0 percentage points while reducing task steps by 12%.

Builders can now plug multiple specialized memory providers into one agent and let MemAgent learn when to retrieve, inject, and store, instead of hand picking a single memory design.

~12 min read← Back to papers

Related papers

Benchmark

According to Me: Long-Term Personalized Referential Memory QA

Jingbiao Mei, Jinghong Chen et al.

arXiv 2026 · 2026

ATM-Bench structurally evaluates long-term multimodal personal memory using Memory Ingestion, Retrieval, and Answer Generation with Schema-Guided Memory and Descriptive Memory variants. On ATM-Bench-Hard, Oracle with SGM reaches 47.3% QS while the best full system stays under 20% accuracy, revealing a large gap.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Questions about this paper?

Paper: MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents

Answers use this explainer on Memory Papers.

Checking…