Paper archive

All AI memory papers

Browse 314 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 3 of 14

BenchmarkAgent MemoryLong-Term Memory

Drawing on Memory: Dual-Trace Encoding Improves Cross-Session Recall in LLM Agents

Benjamin Stern, Peter Nadel

· 2026

Drawing on Memory uses dual-trace memory encoding, an evidence scoring gate, and a three-state retrieval protocol to store paired fact and scene traces in Letta’s archival memory. On LongMemEval-S, Drawing on Memory reaches 73.7% accuracy versus 53.5% for the fact-only C7-control baseline, a +20.2 percentage point gain concentrated in temporal, update, and multi-session questions.

Agent Memory

DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

Sarthak Singh

arXiv 2026 · 2026

DreamBench-SWE wraps a fixed wake agent with a raw episode log, derived memory store, maintenance pipeline, and retrieval gate to test multi-session memory hygiene in software repositories. On 60 traps and 180 S3 cells, the strongest verbatim baseline B5 reaches 89/180 Pass@1 while the hybrid reference probe reaches 95/180, but the clustered P1 comparison fails to reject (p=0.518).

Long-Term Memory

EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents

Xinze Li, Ziyue Zhu et al.

· 2026

EMemBench builds on Interaction episode, Benchmark generator, Structured logging, State reconstruction, and Programmatic QA generation to turn each agent’s own game trajectory into verifiable episodic-memory questions. On text-only games, EMemBench shows A-MEM with Qwen3-32B achieves 51.9% Overall ACC compared to 44.9% for the Qwen3-32B in-context baseline (+7.0 points).

Agent Memory

E-mem: Multi-agent based Episodic Context Reconstruction for LLM Agent Memory

Kaixiang Wang, Yidan Lin et al.

· 2026

E-mem combines a Master Agent, multiple Assistant Agents, a Multi-Pathway Routing Mechanism, Sliding Window Segmentation with Overlap, and Episodic Memory Context Retention and Isolation to reconstruct native contexts instead of compressing them. On the LoCoMo benchmark, E-mem achieves 54.17 F1 with GPT-4o-mini, beating GAM by 8.86 F1 while reducing normalized token cost from 169100 to 3621.

Long-Term Memory

ER-MIA: Black-Box Adversarial Memory Injection Attacks on Long-Term Memory-Augmented Large Language Models

Mitchell Piehl, Zhaohan Xi et al.

· 2026

ER-MIA combines Content-Based MIAs, Question-Targeted MIAs, and an attack Arsenal of primitive and ensemble strategies to poison similarity-based retrieval in long-term memory systems. On the LoCoMo benchmark, ER-MIA’s ensembles cut Mem0’s overall F1 from 23.60% to 2.87% (-87.6%) and reduce single attacks like Harsh Instruction by up to -71.5% F1.

Agent Memory

EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective

Yuyao Wang, Zhongjian Zhang et al.

arXiv 2026 · 2026

EvoMemBench evaluates self-evolving agent memory using four settings and six datasets spanning in-episode and cross-episode, knowledge and execution evolution. EvoMemBench’s main result is that explicit memory only helps reliably when context is insufficient, tasks are hard, or stored experience matches the decision process, while Gemini-3-Flash and other long-context baselines often match or beat 15 memory methods.

BenchmarkAgent Memory

Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents

Xing Zhang, Guanghui Wang et al.

· 2026

Experience Compression Spectrum organizes Level 0 Raw Trace, Level 1 Episodic Memory, Level 2 Procedural Skill, and Level 3 Declarative Rule into a unified scaffold-level compression framework. Experience Compression Spectrum’s mapping of 20+ systems and <1% cross-citation rate shows that all existing agents fix a single compression level and never perform adaptive cross-level compression.

Benchmark

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

Ben Wang, Kang Zhou et al.

arXiv 2026 · 2026

FinPerMA combines Persona Synthesis, Event Timeline Construction, the three-layer Impact Model, Dialogue Synthesis, and Evaluation Task Design to generate frozen, event-grounded investor trajectories for memory benchmarking. On the FinPerMA v8gold corpus, full-context Qwen-3.8 reaches 46.9% overall accuracy and 38.7% MCQ, while retrieval-based memory on Qwen3.7-Max recovers ≈88% of the no-memory–full-context gap with only 1.40k tokens of context.

Agent Memory

ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction

Linhao Zhong, Zongze Du et al.

arXiv 2026 · 2026

ForeDreamer separates factual memory from experiential memory, using a main agent, a memory-processing subagent, an Experience Bank, and a MemGuide–MemTools workspace to process web evidence before forecasting. On Prophet Arena, ForeDreamer achieves an average Brier score of 0.1471 with Qwen3.5-Flash, improving over the Full Text baseline at 0.2059.

Long-Term Memory

From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

Md Nayem Uddin, Kumar Shubham et al.

· 2026

Memora combines Seed Data Design, Session Simulation, Conversation Generation, and Questions and Evaluation Criteria to build weeks-to-months multi-session benchmarks with evolving memories. Using FAMA on Memora, long-term memory agents like MemoBase fall from 43.60 to 15.18 remembering FAMA between weekly and quarterly settings, revealing severe brittleness under heavy memory mutation.

Benchmark

From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms

Jinghao Luo, Yuchen Tian et al.

arXiv 2026 · 2026

From Storage to Experience formalizes LLM agent memory into Storage, Reflection, and Experience, detailing mechanisms like Linear, Vector, Structured storage and Introspection, Environment, Coordination reflection. From Storage to Experience’s main result is a unified evolutionary framework that connects fragmented operating system style and cognitive science inspired memory designs into a single roadmap.

BenchmarkAgent MemoryMemory Architecture

GAM: Hierarchical Graph-based Agentic Memory for LLM Agents

Zhaofen Wu, Hanrong Zhang et al.

· 2026

GAM builds a Hierarchical Graph Memory Architecture with a global Topic Associative Network, local Event Progression Graphs, State-Based Memory Consolidation, and Graph-Guided Multi-Factor Retrieval to decouple encoding from consolidation. On LoCoMo with Qwen2.5-7B, GAM attains an Average F1 of 40.00 compared to Mem0’s 35.38, and on LongDialQA with Qwen2.5-7B, GAM reaches 12.55 F1 vs MemoryOS at 6.76.

BenchmarkBenchmarkAgent MemoryLong-Term Memory

Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework

Chingkwun Lam, Jiaxin Li et al.

· 2026

SSGM interposes a Governance Middleware, Read Filtering Gate, Write Validation Gate, and a dual substrate of Mutable Active Graph plus Immutable Episodic Log between agents and memory. SSGM unifies evolving-memory systems into a four-dimensional failure taxonomy and proves that periodic reconciliation can bound semantic drift over infinite horizons.

Benchmark

GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent

Yuri Kuratov, Matvey Kairov et al.

· 2026

GradMem combines a WRITE phase, a READ phase, a context encoder Eθ, a self-supervised WRITE objective Lwrite, and a meta-learned initialization M0 to optimize prefix memory tokens via test-time gradient descent while keeping model weights frozen. On associative KV-retrieval with 96 key–value pairs, GradMem with 5 gradient WRITE steps reaches 88.4% exact match versus 12.9% for forward-only RMT with the same 8-vector memory.

Agent Memory

Graph-based Agent Memory: Taxonomy, Techniques, and Applications

Chang Yang, Chuang Zhou et al.

· 2026

Graph-based Agent Memory organizes agent memory into Knowledge vs Experience Memory, Short-term vs Long-term Memory, and Non-structural vs Structural Memory tied to graph implementations. It unifies memory extraction, storage, retrieval, and evolution into a single lifecycle, mapping over 70 named systems like MemGPT, GraphRAG, and Zep into a coherent design space.

Benchmark

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Jingbo Yang, Kwei-Herng Lai et al.

arXiv 2026 · 2026

GroupMemBench combines Graph-grounded Message Synthesis, Path-Based Context Sampling, Message Generation, Adversarial Question Synthesis, and a Solve–Judge–Refine Loop to stress-test group memory in LLM agents. On GroupMemBench, Hindsight reaches 46.01% average accuracy while a simple BM25 baseline attains 43.22%, revealing that current memory ingestion pipelines fail to preserve crucial multi-user structure.

Agent Memory

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

Wei-Chieh Huang, Weizhi Zhang et al.

arXiv 2026 · 2026

Harness the Memory instruments a Unified Evaluation Harness, External Memory Substrate Families, Internal Memory Substrate Families, Retrieval Depth Sweep, and Scalability Study to compare 11 memory substrates under identical agents. Harness the Memory shows, for example, that M7 Distilled Strategies reaches 32.1% TSR on ALFWorld-unseen with QWEN3-32B-AWQ, a +9.7 percentage point gain over the NoMem baseline at only 1.23× latency.

RAGLong-Term Memory

HingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable Dialogues

Yijie Zhong, Yunfan Gao, Haofen Wang

· 2026

HingeMem combines Boundary Guided Long-Term Memory, Dialogue Boundary Extraction, Memory Construction, Query Adaptive Retrieval, Hyperedge Rerank, and Adaptive Stop to segment dialogues into element-indexed hyperedges and plan query-specific retrieval. On LOCOMO, HingeMem achieves 63.9 overall F1 and 75.1 LLM-as-a-Judge score, surpassing the best baseline Zep (56.9 F1) by 7.0 F1 without using category-specific QA formats.

Cognitive ArchitectureLong-Term Memory

Human-Like Lifelong Memory: A Neuroscience-Grounded Architecture for Infinite Interaction

Diego C. Lerma-Torres

· 2026

Human-Like Lifelong Memory combines Executive Function and Working Memory, a Memory Service Knowledge Graph, and a Thalamic Gateway to implement dual-process, valence-aware lifelong memory. Human-Like Lifelong Memory is a theoretical framework with seven functional properties and testable predictions rather than benchmark numbers against specific baselines.

Agent Memory

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Ruizhe Li, Mingxuan Du et al.

arXiv 2026 · 2026

Keep It InMind introduces the InMind benchmark plus paired controls (Naive query, Indirect query, Target recall, Backbone control) to isolate failures in agent memory use. On InMind, Keep It InMind finds that the backbone reaches 84.0% indirect accuracy with the memory in context, while six retrieval-based systems reach at most 14.4% application accuracy.