LeanMem: Simple and Efficient Long-Term Memory for LLM Agents

AuthorsYuxin Liao, Le Wu, Min Hou et al.

arXiv 20262026

TL;DR

LeanMem uses controlled routing into profile, event, and record memories plus query-adaptive retrieval to reach 91.80% Accuracy on LongMemEval-S, +15.07 points over A-Mem.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

Uniform Memory Pipelines Waste Tokens and Lose Evidence

Existing long-term memory systems apply a uniform summarization and retrieval pipeline, causing either excessive token consumption or irreversible loss of fine-grained evidence.

When conversational agents compress all history the same way, they discard detail-intensive records and mis-handle evolving events, leading to wrong answers on long-range conversational QA tasks.

HOW IT WORKS

LeanMem: Controlled Memory Writing + Selective Evolution + Adaptive Evidence Composition

Controlled Memory Writing filters key utterances, performs topic segmentation, and uses a scheduler to route segments into profile, event, or record memory with schema-constrained materialization.

Think of LeanMem like a library: profile memory is a compact card catalog of user facts, event memory is a timeline logbook, and record memory is a set of bookmarks into full conversation transcripts.

By combining Selective Memory Evolution for events with Adaptive Evidence Composition at query time, LeanMem preserves temporal structure and detailed sources that a plain context window or single-summary memory cannot.

DIAGRAM

Query-Time Adaptive Evidence Composition Flow

This diagram shows how LeanMem plans retrieval and composes evidence from profile, event, and record memories for a given query.

DIAGRAM

LeanMem Evaluation and Ablation Pipeline

This diagram shows how LeanMem is evaluated on LoCoMo and LongMemEval-S, including ablations of its main components.

PROCESS

How LeanMem Handles a Long-Term Conversational Question

  1. 01

    Controlled Memory Writing

    LeanMem filters key utterances, applies topic segmentation, and uses write scheduling to route each segment into profile, event, or record memory with schema-constrained materialization.

  2. 02

    Memory Materialization

    LeanMem stores extracted attribute pairs as profile memory, temporal anchors and states as event memory, and retrieval gists plus GLiNER keywords and source indices as record memory.

  3. 03

    Selective Memory Evolution

    LeanMem buffers new event memories, periodically retrieves related events by topic, orders them by time, and incrementally merges them into updated temporal state representations.

  4. 04

    Adaptive Evidence Composition

    Given a query, LeanMem builds a retrieval plan πq, selects memory types, allocates type-specific top-k budgets, reranks candidates with constraints Cq, and assembles ordered evidence for answer generation.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Controlled Memory Writing

    LeanMem introduces Controlled Memory Writing with key utterance filtering, topic segmentation, and write scheduling to route segments into profile, event, or record memory, cutting LoCoMo build tokens to 69.45K.

  • 02

    Selective Memory Evolution

    LeanMem performs Selective Memory Evolution only on event memory via deferred, localized consolidation, improving LongMemEval-S Accuracy to 91.80% while keeping construction tokens at 117.61K.

  • 03

    Adaptive Evidence Composition

    LeanMem proposes Adaptive Evidence Composition that plans retrieval across memory types, achieving 97.67 Recall and 91.80 Accuracy on LongMemEval-S with GPT-4.1-mini, +15.07 Accuracy points over A-Mem.

RESULTS

By the Numbers

Accuracy

91.80%

+15.07 over A-Mem on LongMemEval-S (71.40%) with GPT-4.1-mini

Recall

97.67%

+2.52 over SimpleMem on LongMemEval-S (95.15%) with GPT-4.1-mini

Build Tokens

117.61K

vs 1330.77K for A-Mem on LongMemEval-S with GPT-4.1-mini

Latency

2.16s

average inference latency per question on LongMemEval-S with GPT-4.1-mini

Table 1 reports LeanMem on LongMemEval-S, a 500-history benchmark averaging 115K tokens and 40–50 sessions per user. The 91.80% Accuracy and 97.67% Recall show that LeanMem preserves long-range evidence while drastically reducing construction tokens compared to A-Mem.

BENCHMARK

By the Numbers

Table 1 reports LeanMem on LongMemEval-S, a 500-history benchmark averaging 115K tokens and 40–50 sessions per user. The 91.80% Accuracy and 97.67% Recall show that LeanMem preserves long-range evidence while drastically reducing construction tokens compared to A-Mem.

BENCHMARK

Overall Effectiveness on LongMemEval-S (GPT-4.1-mini)

Accuracy on LongMemEval-S conversational QA with different memory systems.

BENCHMARK

Construction Cost on LongMemEval-S (GPT-4.1-mini)

Build tokens per conversation for different memory systems.

KEY INSIGHT

The Counterintuitive Finding

LeanMem reaches 91.80% Accuracy on LongMemEval-S while using only 117.61K construction tokens, compared to A-Mem’s 71.40% Accuracy with 1330.77K tokens.

This is surprising because we usually expect higher token budgets and more aggressive summarization to help accuracy, yet LeanMem’s heterogeneous, loss-aware memory actually needs far fewer tokens.

WHY IT MATTERS

What this unlocks for the field

LeanMem makes it practical to run LLM agents over histories averaging 115K tokens and up to 50 sessions without blowing context windows or token budgets.

Builders can now design agents that track evolving user states, preserve detailed plans, and answer long-range questions reliably using profile, event, and record memories instead of monolithic summaries.

~12 min read← Back to papers

Related papers

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

Questions about this paper?

Paper: LeanMem: Simple and Efficient Long-Term Memory for LLM Agents

Answers use this explainer on Memory Papers.

Checking…