Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 5 of 19

BenchmarkAgent MemoryLong-Term Memory

Drawing on Memory: Dual-Trace Encoding Improves Cross-Session Recall in LLM Agents

Benjamin Stern, Peter Nadel

· 2026

Drawing on Memory uses dual-trace memory encoding, an evidence scoring gate, and a three-state retrieval protocol to store paired fact and scene traces in Letta’s archival memory. On LongMemEval-S, Drawing on Memory reaches 73.7% accuracy versus 53.5% for the fact-only C7-control baseline, a +20.2 percentage point gain concentrated in temporal, update, and multi-session questions.

Agent Memory

DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

Sarthak Singh

arXiv 2026 · 2026

DreamBench-SWE wraps a fixed wake agent with a raw episode log, derived memory store, maintenance pipeline, and retrieval gate to test multi-session memory hygiene in software repositories. On 60 traps and 180 S3 cells, the strongest verbatim baseline B5 reaches 89/180 Pass@1 while the hybrid reference probe reaches 95/180, but the clustered P1 comparison fails to reject (p=0.518).

Agent Memory

Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation

Wenzhi Li, Dong Nie et al.

arXiv 2026 · 2026

Dual-Layer Agentic Memory combines a fast write router, small-to-large cost-aware routing cascade, operational memory taxonomy, and write-back consolidation to manage parametric and external memory. On the streaming ZsRE benchmark, Dual-Layer Agentic Memory’s Write Router retains 86.35% QA EM versus 87.79% for Full Store, while cutting storage to 77.00% and reducing 8B routing compute by up to 39.7%.

Benchmark

DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings

Wenya Xie, Shengming Zhou et al.

arXiv 2026 · 2026

DynamicMem builds 15‑month, multi‑app trajectories using Multi-timescale user profile construction, Intent-conditioned event chain generation, and State-consistent multi-app log generation to stress-test long-horizon memory. DynamicMem’s main result shows State Completion declines by up to 26.5 points for A-Mem from checkpoint C1 to C5, while Personalized Service scores remain stable or improve, revealing hidden failure modes in memory systems.

Memory Architecture

ECHO: A Cognitively Inspired, Auditable Memory Plane for Long-Horizon Agents

Yu Qian, Hong Miao et al.

arXiv 2026 · 2026

ECHO orchestrates Immutable events, Typed projection, Bitemporal ledger, Candidate discovery, Provenance closure, and Derive then realize into a governed memory plane for long-horizon agents. On LongMemEval-S, ECHO reaches 97.60% Hit@10 and 88.84% turn Recall@5, while a matched QA sample shows Mem0 OSS at 64.84% versus ECHO’s 41.76% (−23.08 percentage points).

Memory Architecture

EchoPath: Execution-Level Replayable Memory for GUI Agents

Yao Zhao, Aditya Shanmugham et al.

arXiv 2026 · 2026

EchoPath wraps GUI agents with ActionLens, an EchoPath Memory Repository, a Replay Module, Memory Consolidation, and an Image-based Target-Reaiming (IBTR) algorithm to turn validated trajectories into callable memories. On OSWorld-Verified tasks, EchoPath achieves 91.2% second-pass success with codex-gpt-5.5-medium while reducing median token usage from about 572k to 20,370 compared to Synapse’s 586,386.

Agent Memory

EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph

Zeyang Cui, Jiannong Cao et al.

arXiv 2026 · 2026

EdgeMem stores verbatim interaction turns and organizes them via a Multi-anchor Hypergraph Construction, Time Sub-Hypergraph, Co-occurrence Sub-Hypergraph, and Episode Sub-Hypergraph with deterministic retrieval. On LoCoMo, EdgeMem reaches 61.01 strict-judge accuracy versus 58.70 for CompassMem while requiring 0 construction tokens and only 1,000 tokens per question.

Benchmark

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

Yijun Chen, Yaqi Zheng et al.

arXiv 2026 · 2026

EM2Mem organizes long videos into Event-Centric Multimodal Memory Cells, Temporal Context Views, and event-linked Episodic and Semantic Graphs for align-then-retrieve reasoning. On Video-MME (L), EM2Mem reaches 76.8% average accuracy, beating the strongest memory baseline WorldMM† at 73.1% and cutting total inference tokens by 63.66%.

Long-Term Memory

EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents

Xinze Li, Ziyue Zhu et al.

· 2026

EMemBench builds on Interaction episode, Benchmark generator, Structured logging, State reconstruction, and Programmatic QA generation to turn each agent’s own game trajectory into verifiable episodic-memory questions. On text-only games, EMemBench shows A-MEM with Qwen3-32B achieves 51.9% Overall ACC compared to 44.9% for the Qwen3-32B in-context baseline (+7.0 points).

Agent Memory

E-mem: Multi-agent based Episodic Context Reconstruction for LLM Agent Memory

Kaixiang Wang, Yidan Lin et al.

· 2026

E-mem combines a Master Agent, multiple Assistant Agents, a Multi-Pathway Routing Mechanism, Sliding Window Segmentation with Overlap, and Episodic Memory Context Retention and Isolation to reconstruct native contexts instead of compressing them. On the LoCoMo benchmark, E-mem achieves 54.17 F1 with GPT-4o-mini, beating GAM by 8.86 F1 while reducing normalized token cost from 169100 to 3621.

Memory Architecture

ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents

Xing Fu, Yulin Hu et al.

arXiv 2026 · 2026

ENPMR-Bench evaluates Emotional Need-aware Proactive Memory Retrieval by tying User Profile Generation, Structured Memory Retrieval Guidelines, Emotional Support Dialogue Generation, and Memory Retrieval Task into a controlled benchmark. On ENPMR-Bench, Qwen3-Embedding-8B reaches only 46.42% Recall@10, while DeepSeek-V3 empathy rises from 4.38 to 4.91 when ENPMR-Bench supplies golden memories instead of retrieved ones.

Agent Memory

EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory

Xuanyu Meng, Xing Fan et al.

arXiv 2026 · 2026

EnSIMem organizes long dialogues via Theme-coherent episodic memory construction, Dialogue-grounded entity-property indexing, Granularity-controlled property alignment, and Requirement-aware evidence localization into an entity structured index linked to original episodes. On LongMemEval, EnSIMem reaches 92.8% average accuracy versus 87.4% for MEMORA (P), and on LoCoMo EnSIMem reaches 90.6% versus 86.3%.

Agent Memory

EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents

Fengzhou Sun, Yuan Zhang et al.

arXiv 2026 · 2026

EP-Mem combines a User Pre-configured Privacy Policy File, Elastic Privacy Memory Unit, EP-Mem Sidecar Overlay, Evolvable EP-Mem Privacy Engine, and Memory Material Preprocessing to enforce social relationship-aware disclosure across long-term agent memory. On EP-Bench Task 3, EP-Mem+Evo achieves PB=0.136 versus 0.738 for NoEngine_evid while improving MIQ from 3.30 to 4.32.

Long-Term Memory

ER-MIA: Black-Box Adversarial Memory Injection Attacks on Long-Term Memory-Augmented Large Language Models

Mitchell Piehl, Zhaohan Xi et al.

· 2026

ER-MIA combines Content-Based MIAs, Question-Targeted MIAs, and an attack Arsenal of primitive and ensemble strategies to poison similarity-based retrieval in long-term memory systems. On the LoCoMo benchmark, ER-MIA’s ensembles cut Mem0’s overall F1 from 23.60% to 2.87% (-87.6%) and reduce single attacks like Harsh Instruction by up to -71.5% F1.

Agent Memory

ERRAND: Budgeted Maintenance of Agent Memory

Beining Wu, Zihao Ding, Jun Huang

arXiv 2026 · 2026

ERRAND combines drift-aware beliefs, a single-peaked errand index, a deadband gate, and a versioned lifecycle to decide when an agent should spend actions revalidating memory instead of working. On the main scripted tool-use world, ERRAND reaches 71.9% ITT success at b=12 compared to 61.9% for eager revalidation (+10.0pp) while spending only 11.0% of steps when uncapped.

Benchmark

ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support

Tiantian Chen, Jiaqi Lu et al.

· 2026

ES-MemEval evaluates conversational agents’ long-term memory via Question Answering, Summarization, and Dialogue Generation on the EvoEmo dataset, which combines User Profile Construction, Event Timeline Expansion, and Chat Data Generation. On ES-MemEval QA, Mistral-24B + RAG reaches 18.8% overall F1 and 1.27 LLM-as-Judge score, improving over base Mistral-24B and revealing how explicit memory and retrieval reshape performance.

RAG

EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

Zeyu Liu, Jian Zhong et al.

arXiv 2026 · 2026

EvalMem decomposes long-term memory systems into Encoding Examiner, Retrieval Examiner, Generation Examiner, and an Attribution Agent that assigns 11 fine-grained defect codes. EvalMem shows retrieval defects reach 22.1% on LOCOMO, and a MemWiki auxiliary index guided by this diagnosis improves mean accuracy by 2.5 percentage points over seven memory systems.

Agent Memory

EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory

Ye Shen, Dun Pei et al.

arXiv 2026 · 2026

EvolMem combines Topic-Initiated Generation, Narrative-Inspired Transformation, Filtering, and Challenge Injection to build 1,600 multi-session dialogues covering declarative and non-declarative memory. On EvolMem, Gemini-3-Pro achieves 67.47% overall memory performance, ahead of GPT-5.1 at 52.82%, while agents like MemoryOS fail to surpass their DeepSeek-V3.2 base.

Agent Memory

EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective

Yuyao Wang, Zhongjian Zhang et al.

arXiv 2026 · 2026

EvoMemBench evaluates self-evolving agent memory using four settings and six datasets spanning in-episode and cross-episode, knowledge and execution evolution. EvoMemBench’s main result is that explicit memory only helps reliably when context is insufficient, tasks are hard, or stored experience matches the decision process, while Gemini-3-Flash and other long-context baselines often match or beat 15 memory methods.

BenchmarkAgent Memory

Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents

Xing Zhang, Guanghui Wang et al.

· 2026

Experience Compression Spectrum organizes Level 0 Raw Trace, Level 1 Episodic Memory, Level 2 Procedural Skill, and Level 3 Declarative Rule into a unified scaffold-level compression framework. Experience Compression Spectrum’s mapping of 20+ systems and <1% cross-citation rate shows that all existing agents fix a single compression level and never perform adaptive cross-level compression.

Benchmark

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

Ben Wang, Kang Zhou et al.

arXiv 2026 · 2026

FinPerMA combines Persona Synthesis, Event Timeline Construction, the three-layer Impact Model, Dialogue Synthesis, and Evaluation Task Design to generate frozen, event-grounded investor trajectories for memory benchmarking. On the FinPerMA v8gold corpus, full-context Qwen-3.8 reaches 46.9% overall accuracy and 38.7% MCQ, while retrieval-based memory on Qwen3.7-Max recovers ≈88% of the no-memory–full-context gap with only 1.40k tokens of context.