EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

AuthorsZeyu Liu, Jian Zhong, Rongduo Han et al.

arXiv 20262026

TL;DR

EvalMem uses three parallel Examiners plus recall-first agentic RAG to diagnose memory failures, revealing retrieval defects up to 22.1% on LOCOMO and enabling +2.5 pp accuracy via MemWiki.

SharePost on XLinkedIn

Read our summary here, or open the publisher PDF on the next tab.

THE PROBLEM

End-to-end QA hides where memory systems break (retrieval defects reach 22.1%)

Existing long-term memory evaluations only report end-to-end QA correctness, so a wrong answer gives just 1-bit feedback and hides the failure layer.

On LOCOMO, EvalMem shows retrieval defects reach 22.1%, while encoding and generation defects are only 7.7% and 6.5%, meaning systems misattribute many errors and struggle to fix them.

HOW IT WORKS

EvalMem: Encoding, Retrieval, Generation Examiners plus Attribution Agent

EvalMem’s core mechanism is three parallel components: Encoding Examiner, Retrieval Examiner, Generation Examiner, coordinated by an Attribution Agent that outputs 11-code multi-label diagnoses.

You can think of EvalMem like a computer debugger: one probe checks what’s written to disk, another inspects cache hits, and a third tests whether the CPU can compute correctly from perfect inputs.

This operation-level view lets EvalMem pinpoint whether a long-term assistant failed to store, retrieve, or use a fact, something a plain context window or single QA score cannot reveal.

DIAGRAM

EvalMem’s Operation-Level Diagnosis Flow

This diagram shows how EvalMem sequences Encoding, Retrieval, and Generation Examiners plus the Attribution Agent to diagnose each query.

DIAGRAM

EvalMem Evaluation Pipeline across Datasets and Systems

This diagram shows how EvalMem runs seven memory systems on LOCOMO, LONGMEMEVAL-S, and DynaMem-Bench and then aggregates defect statistics.

PROCESS

How EvalMem Handles a Query — Encoding–Retrieval–Generation Lifecycle

  1. 01

    Encoding Examiner

    EvalMem uses the Encoding Examiner to check whether key facts Fkey are stored in the memory store M, assigning states like EXIST, MISS, CORRUPTAmbig, CORRUPTWrong, or DIRTY.

  2. 02

    Retrieval Examiner

    EvalMem’s Retrieval Examiner inspects the native Coriginal, tests whether Fkey is supported, and classifies retrieval as HIT, MISS, or NOISE with extra RF, LATE, NOI, and NIR codes.

  3. 03

    Generation Examiner

    EvalMem’s Generation Examiner feeds Q plus oracle evidence Coracle to the answer model, judging PASS or FAIL and subtyping failures into GF, GRF, or GH depending on reasoning and hallucination.

  4. 04

    Attribution Agent

    EvalMem’s Attribution Agent merges Denc, Dret, and Dgen into an 11-code multi-label defect set, masking downstream retrieval defects when encoding MISS or dirty memory explains the error.

KEY CONTRIBUTIONS

Key Contributions

  • 01

    Fine-grained attribution for memory-system evaluation

    EvalMem introduces three Examiners plus an Attribution Agent that produce an 11-code multi-label diagnosis, separating EM, RF, GF, GRF, GH, and other defects for each query.

  • 02

    Reliable and dynamic diagnostic evaluation

    EvalMem adapts recall-first agentic RAG inside the Encoding Examiner, reducing observation MISS from 29.8% to 4.4% on LOCOMO, and builds DynaMem-Bench with 962 questions over 10 personas.

  • 03

    Retrieval as the most frequently attributed failure layer

    EvalMem shows retrieval defects dominate encoding and generation across datasets, reaching 22.1% vs 7.7% and 6.5% on LOCOMO, and guides MemWiki to improve accuracy by +2.5 pp.

RESULTS

By the Numbers

Mean Accuracy LOCOMO

68.1%

+? over static end-to-end scores (EvalMem adds defect attribution, not a new baseline)

Retrieval Defect Rate LOCOMO

22.1%

+14.4 pp over Encoding defect rate 7.7%

MemWiki Accuracy Gain LOCOMO

+2.5 pp

guided by retrieval diagnosis using EvalMem

Agentic RAG MISS Reduction

29.8% → 4.4%

−25.4 pp observation MISS on LOCOMO POS with present evidence

EvalMem evaluates seven memory systems on LOCOMO, LONGMEMEVAL-S, and DynaMem-Bench, exposing retrieval as the dominant failure layer. The MemWiki intervention and recall-first agentic RAG show that EvalMem’s diagnoses are reliable and actionable for improving long-term memory systems.

BENCHMARK

By the Numbers

EvalMem evaluates seven memory systems on LOCOMO, LONGMEMEVAL-S, and DynaMem-Bench, exposing retrieval as the dominant failure layer. The MemWiki intervention and recall-first agentic RAG show that EvalMem’s diagnoses are reliable and actionable for improving long-term memory systems.

BENCHMARK

Layer-level defect rates on LOCOMO (Set 1: GPT-4O-MINI)

Encoding vs Retrieval vs Generation defect incidence across seven memory systems on LOCOMO.

KEY INSIGHT

The Counterintuitive Finding

EvalMem shows that retrieval defects consistently exceed encoding and generation defects, reaching 22.1% vs 7.7% and 6.5% on LOCOMO and 25.9% vs 9.7% and 5.0% on DynaMem-Bench.

This is counterintuitive because many memory-system designs focus on better storage or summarization, yet EvalMem reveals that the main bottleneck is actually surfacing already stored facts reliably.

WHY IT MATTERS

What this unlocks for the field

EvalMem unlocks operation-level visibility into long-term memory systems, letting builders see exactly whether encoding, retrieval, or generation failed for each query.

With EvalMem, practitioners can now design targeted interventions like MemWiki or retriever changes, validate them quantitatively, and iterate beyond opaque end-to-end QA scores that previously hid these failure modes.

~12 min read← Back to papers

Related papers

RAG

A Dynamic Retrieval-Augmented Generation System with Selective Memory and Remembrance

Okan Bursa

· 2026

Adaptive RAG Memory (ARM) augments a standard retriever–generator stack with a Dynamic Embedding Layer and Remembrance Engine that track usage statistics and apply selective remembrance and decay to embeddings. On a lightweight retrieval benchmark, ARM achieves NDCG@5 ≈ 0.9401 and Recall@5 = 1.000 with 22M parameters, matching larger baselines like gte-small while providing the best efficiency among ultra-efficient models.

RAG

Are We Ready For An Agent-Native Memory System?

Wei Zhou, Xuanhe Zhou et al.

arXiv 2026 · 2026

Are We Ready For An Agent-Native Memory System? analyzes Memory Representation and Storage, Memory Extraction, Memory Retrieval and Routing, and Memory Maintenance across 12 real systems like Mem0, Zep, MemTree, LightMem, MemOS, MemoryOS, and A-MEM. The study’s main result is that structured systems such as Zep reach 48.0 LLM Judge Accuracy on LongMemEval while Long Context reaches 19.0, and that localized maintenance strategies like LightMem achieve 48.3% normalized utility at only 3.67 s per query.

Questions about this paper?

Paper: EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

Answers use this explainer on Memory Papers.

Checking…