Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

AuthorsYuanyi Song, Yukai Wang, Xinbei Ma et al.

arXiv 20262026

TL;DR

REALM uses retrieval-driven memory reconsolidation over a heterogeneous cognitive graph to reach 75.97% accuracy on LoCoMo, +7.17 points over MAGMA.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Static long term memories ignore retrieval feedback

Existing long term memory systems mainly update memories with new information using predefined data structures and a fixed retrieval pipeline.

When retrieval is treated as a passive endpoint, agents miss feedback driven reorganization, leading to brittle memory topology and weaker future reasoning.

HOW IT WORKS

REALM — Retrieval driven memory reconsolidation lifecycle

REALM combines Unified Cognitive Graph, Memory Construction, Strategy Atom Combination, and Local Topology Reconsolidation into a closed loop memory lifecycle.

Think of REALM as a brain like card catalog: retrieval not only looks up entries but rewires the graph, strengthening or weakening connections after each use.

This retrieval driven reconsolidation lets REALM evolve beyond a static context window, forming usage aware structures that cluster evidence for multi hop and temporal reasoning.

DIAGRAM

REALM query time retrieval and reconsolidation flow

This diagram shows how REALM retrieves evidence via strategy atoms and then reconsolidates the activated subgraph after each query.

DIAGRAM

REALM evaluation pipeline on LoCoMo and LongMemEval

This diagram shows how REALM is evaluated across LoCoMo and LongMemEval with GPT-4o-mini as backbone and judge.

PROCESS

How REALM Handles a Long horizon interaction lifecycle

  1. 01

    Memory Organization

    REALM builds a Unified Cognitive Graph where nodes and edges encode entities, events, episodes, facts, and their relations for later retrieval.

  2. 02

    Memory Construction

    During integration, REALM’s Memory Construction phase decides to add, merge, or skip candidate units and infers relational edges to update the graph.

  3. 03

    Memory Retrieval

    Using Strategy Atom Combination, REALM composes seed, expansion, and filtering policies to traverse the cognitive graph and collect evidence Vq.

  4. 04

    Memory Reconsolidation

    In Local Topology Reconsolidation, REALM applies add, strengthen, or weaken decisions on edges in the activated subgraph Gq to evolve future retrieval paths.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Memory Lifecycle Concept

    REALM formalizes long term memory as a closed loop lifecycle over a Unified Cognitive Graph, where retrieval triggers structural updates instead of leaving memory static.

  • 02

    REALM Framework

    REALM integrates Memory Construction, Strategy Atom Combination, and Local Topology Reconsolidation into one agentic framework for autonomous organization and adaptive retrieval.

  • 03

    Empirical Validation

    REALM reaches 75.97% accuracy on LoCoMo and 65.11% on LongMemEval, with reconsolidation adding up to +6.66 points on single session preference questions.

RESULTS

By the Numbers

LoCoMo Average

75.97%

+7.17 over MAGMA

LongMemEval Average

65.11%

+1.31 over Zep

LoCoMo Multi Hop

64.54%

+11.74 over Nemori

LongMemEval Knowledge update

88.89%

+14.49 over MAGMA

REALM is evaluated on LoCoMo and LongMemEval_S, which test multi hop, temporal, open domain, and multi session memory use. These results show REALM’s retrieval driven reconsolidation yields large gains over MAGMA and Zep on challenging long term reasoning tasks.

BENCHMARK

By the Numbers

REALM is evaluated on LoCoMo and LongMemEval_S, which test multi hop, temporal, open domain, and multi session memory use. These results show REALM’s retrieval driven reconsolidation yields large gains over MAGMA and Zep on challenging long term reasoning tasks.

BENCHMARK

Accuracy on LoCoMo by method

Average accuracy (%) on LoCoMo using GPT-4o-mini as backbone and judge.

BENCHMARK

Accuracy on LongMemEval by method

Average accuracy (%) on LongMemEval_S using GPT-4o-mini as backbone and judge.

KEY INSIGHT

The Counterintuitive Finding

Enabling reconsolidation increases correct answers by up to +6.66 points on single session preference, while evidence discovery rises by only 3.70% and 5.41%.

This is surprising because we expect better performance mainly from retrieving more memories, but REALM shows reorganizing existing evidence paths matters far more.

WHY IT MATTERS

What this unlocks for the field

REALM unlocks long term agents whose memory graphs adapt through usage, strengthening successful reasoning paths and weakening misleading ones automatically.

Builders can now design agents that learn their own retrieval topology over time, moving beyond fixed memory structures and static pipelines that ignore retrieval feedback.

~12 min read← Back to papers

Related papers

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

Questions about this paper?

Paper: Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

Answers use this explainer on Memory Papers.

Checking…

Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents | Memory Papers