Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents

AuthorsWanqi Zhou, Jiawei Lu, Yang Wang et al.

arXiv 20262026

TL;DR

RIME uses question-guided retrieval-induced memory evolution to reach 84.29 Judge on LoCoMo with GPT-5.6 Sol, beating Nemori by +1.24 while using far fewer tokens.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Compression-first memory loses important dialogue evidence

Existing long-term memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression.

When future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that later matter for answering user queries.

HOW IT WORKS

RIME: Retrieval-induced memory evolution

RIME introduces Question-guided evidence recall, Historical memory grounding, Joint memory formation and reconciliation, and Selective source-context recovery to evolve a compact yet recoverable memory bank.

You can think of RIME like a librarian using fixed question cards to pull relevant books and notes before rewriting the catalog, instead of compressing the whole archive blindly.

This retrieval-induced evolution lets RIME preserve and later recover fine-grained dialogue evidence that a plain context window or one-shot compression would discard.

DIAGRAM

Selective memory use during answering

This diagram shows how RIME answers a user query by first using semantic memories and then selectively retrieving source dialogue when needed.

DIAGRAM

LoCoMo evaluation and ablation pipeline

This diagram shows how RIME is evaluated on LoCoMo, including baselines, backbones, and ablation variants.

PROCESS

How RIME Handles a Dialogue Session

  1. 01

    Question-guided evidence recall

    RIME uses a fixed set of six generic formation questions to retrieve kf relevant dialogue turns from the current session, surfacing focused evidence before compression.

  2. 02

    Local context expansion

    RIME expands recalled turns with two rounds of neighboring context using Nt, building an evidence set Et that preserves dependencies and surrounding dialogue.

  3. 03

    Historical memory grounding

    RIME concatenates each question with its recalled evidence to form zt and retrieves kh relevant historical memories Ht from Mt−1 for grounding and conflict detection.

  4. 04

    Joint memory formation and reconciliation

    RIME calls Fmem over Et and Ht to predict ADD, UPDATE, DELETE, and NOOP operations, then applies them to update the memory bank Mt with provenance and timestamps.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Question-guided retrieve-then-consolidate mechanism

    RIME introduces Question-guided evidence recall and Historical memory grounding so generic questions retrieve focused dialogue evidence before consolidation, improving Judge from 71.62 to 81.82 over direct compression.

  • 02

    Selective source-context recovery for memory use

    RIME adds Selective source-context recovery that retrieves only query-relevant source turns when semantic memory fails, boosting Judge by 10.59 points on Qwen3-235B-A22B compared to disabling this mechanism.

  • 03

    Efficient long-term memory with fewer tokens

    RIME achieves the strongest Judge, F1, and BLEU-1 on LoCoMo while using 5.11K tokens with Qwen3-235B-A22B, versus 63.19K for A-MEM and 20.92K for Nemori.

RESULTS

By the Numbers

Judge

84.29

+1.24 over Nemori with GPT-5.6 Sol

F1

56.41

+3.77 over A-MEM with GPT-5.6 Sol

BLEU-1

48.85

+2.97 over Nemori with GPT-5.6 Sol

Tokens

5.21K

−12.53K vs Nemori and −46.93K vs A-MEM with GPT-5.6 Sol

On the 1,540 non-adversarial LoCoMo questions, which test single-hop, multi-hop, temporal, and open-domain memory, RIME delivers the best Judge, F1, and BLEU-1 across all compared systems. These results show that RIME’s retrieval-induced memory evolution yields both higher answer quality and dramatically lower query-time LLM token usage than deliberative memory baselines.

BENCHMARK

By the Numbers

On the 1,540 non-adversarial LoCoMo questions, which test single-hop, multi-hop, temporal, and open-domain memory, RIME delivers the best Judge, F1, and BLEU-1 across all compared systems. These results show that RIME’s retrieval-induced memory evolution yields both higher answer quality and dramatically lower query-time LLM token usage than deliberative memory baselines.

BENCHMARK

Main results on LoCoMo with GPT-5.6 Sol

Judge score on 1,540 non-adversarial LoCoMo questions.

BENCHMARK

Effect of memory formation strategies on LoCoMo (Qwen3-235B-A22B)

Judge score for different memory formation strategies with the same selective memory-use stage.

KEY INSIGHT

The Counterintuitive Finding

RIME reaches a Judge score of 81.82 with Qwen3-235B-A22B using question-guided formation, compared to 71.62 for Direct Compression with the same backbone.

This is surprising because simply adding question prompts to compression (72.66 Judge) barely helps, yet turning those questions into retrieval cues yields a +10.20 Judge jump.

WHY IT MATTERS

What this unlocks for the field

RIME unlocks long-term agents that can both maintain compact semantic memories and selectively recover omitted dialogue context without scanning full histories.

Builders can now design LLM agents that stay coherent over many sessions while keeping inference costs manageable, using retrieval-induced memory evolution instead of monolithic compression or expensive deliberation.

~12 min read← Back to papers

Related papers

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

Questions about this paper?

Paper: Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents

Answers use this explainer on Memory Papers.

Checking…