Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 6 of 19

Memory Architecture

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Minghao Li, Bangyan Li et al.

arXiv 2026 · 2026

FlowState represents agent memory as persistent execution state using a Persistent State Repository, a Historical State Index, and an Active State coordinated by Incremental State Update and Progressive State Access. On MemoryArena, FlowState improves average success rate over Full Context with DeepSeek-V4-Flash by 4.55 percentage points while reducing total token consumption by 43.2%.

Agent Memory

ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction

Linhao Zhong, Zongze Du et al.

arXiv 2026 · 2026

ForeDreamer separates factual memory from experiential memory, using a main agent, a memory-processing subagent, an Experience Bank, and a MemGuide–MemTools workspace to process web evidence before forecasting. On Prophet Arena, ForeDreamer achieves an average Brier score of 0.1471 with Qwen3.5-Flash, improving over the Full Text baseline at 0.2059.

Agent Memory

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

Evan Chen, Shiqiang Wang, Christopher G. Brinton

arXiv 2026 · 2026

PLANFENCE attaches exact parent IDs to plans, validates declared dependencies with authoritative owners, and enforces a dependency-scoped action gate using components like Exact derivation, Declared scope, and Dependency-scoped action gate. In the high-churn AT&T setting, PLANFENCE achieves 0/330 invalid actions with 230.8 ms stall and 8.1 KiB traffic per action, whereas Local replica and Owner-head freshness each issue 330/330 invalid actions.

Long-Term Memory

From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

Md Nayem Uddin, Kumar Shubham et al.

· 2026

Memora combines Seed Data Design, Session Simulation, Conversation Generation, and Questions and Evaluation Criteria to build weeks-to-months multi-session benchmarks with evolving memories. Using FAMA on Memora, long-term memory agents like MemoBase fall from 43.60 to 15.18 remembering FAMA between weekly and quarterly settings, revealing severe brittleness under heavy memory mutation.

Benchmark

From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms

Jinghao Luo, Yuchen Tian et al.

arXiv 2026 · 2026

From Storage to Experience formalizes LLM agent memory into Storage, Reflection, and Experience, detailing mechanisms like Linear, Vector, Structured storage and Introspection, Environment, Coordination reflection. From Storage to Experience’s main result is a unified evolutionary framework that connects fragmented operating system style and cognitive science inspired memory designs into a single roadmap.

BenchmarkAgent MemoryMemory Architecture

GAM: Hierarchical Graph-based Agentic Memory for LLM Agents

Zhaofen Wu, Hanrong Zhang et al.

· 2026

GAM builds a Hierarchical Graph Memory Architecture with a global Topic Associative Network, local Event Progression Graphs, State-Based Memory Consolidation, and Graph-Guided Multi-Factor Retrieval to decouple encoding from consolidation. On LoCoMo with Qwen2.5-7B, GAM attains an Average F1 of 40.00 compared to Mem0’s 35.38, and on LongDialQA with Qwen2.5-7B, GAM reaches 12.55 F1 vs MemoryOS at 6.76.

Benchmark

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

Zhe Ren, Yibo Yang et al.

arXiv 2026 · 2026

GateMem structures shared-memory evaluation around Episodes and Memory State, Checkpoints and Governance Categories, and a multiplicative Memory Governance Score that couples utility, access violations, and forgetting failures. GateMem shows, for example, GPT-5.4 with LONG-CONTEXT reaches 80.1% MGS on the Medical domain while RAG-NAIVE only achieves 44.7%, yet even LONG-CONTEXT still leaks unauthorized or deleted information.

BenchmarkBenchmarkAgent MemoryLong-Term Memory

Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework

Chingkwun Lam, Jiaxin Li et al.

· 2026

SSGM interposes a Governance Middleware, Read Filtering Gate, Write Validation Gate, and a dual substrate of Mutable Active Graph plus Immutable Episodic Log between agents and memory. SSGM unifies evolving-memory systems into a four-dimensional failure taxonomy and proves that periodic reconciliation can bound semantic drift over infinite horizons.

Benchmark

GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent

Yuri Kuratov, Matvey Kairov et al.

· 2026

GradMem combines a WRITE phase, a READ phase, a context encoder Eθ, a self-supervised WRITE objective Lwrite, and a meta-learned initialization M0 to optimize prefix memory tokens via test-time gradient descent while keeping model weights frozen. On associative KV-retrieval with 96 key–value pairs, GradMem with 5 gradient WRITE steps reaches 88.4% exact match versus 12.9% for forward-only RMT with the same 8-vector memory.

Agent Memory

Graph-based Agent Memory: Taxonomy, Techniques, and Applications

Chang Yang, Chuang Zhou et al.

· 2026

Graph-based Agent Memory organizes agent memory into Knowledge vs Experience Memory, Short-term vs Long-term Memory, and Non-structural vs Structural Memory tied to graph implementations. It unifies memory extraction, storage, retrieval, and evolution into a single lifecycle, mapping over 70 named systems like MemGPT, GraphRAG, and Zep into a coherent design space.

Benchmark

Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation

Dac Duy Anh Nguyen, Zhangchi Qiu et al.

arXiv 2026 · 2026

Graph-Based Personalized Memory for LLM Agents organizes memory representation, memory evolution, memory retrieval, and memory evaluation into a single lifecycle for LLM personalization. This unified taxonomy clarifies how graph-based memory supports long-horizon personalized assistants across benchmarks like LoCoMo, PersonaMem-v2, and EngramaBench.

Benchmark

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

Geng Li, Yuhao Wang et al.

arXiv 2026 · 2026

GraphMemix combines candidate graph construction, evidence utility and activation costs, and forest optimization to organize long-term multimodal memories as a query-conditioned evidence forest. On four personal multimodal memory benchmarks with Qwen3-VL-8B-Instruct, GraphMemix reaches 61.55% macro Judge Accuracy on ATM, Mem-Gallery, MemEye, and H2HMem, improving over UniversalRAG’s 49.80% by 11.75 percentage points.

Memory Architecture

Grounding Memory Summarization in Utility Intent

Zhenyu Lei, Mingjia Shi et al.

arXiv 2026 · 2026

MemSuit combines Adaptive Entry Decomposition, Utility-Aware Self-Distillation, and Entry-Aware Retrieval Calibration to turn raw conversations into compact, query-useful memory entries and aligned embeddings. On LoCoMo with Llama-3.1-8B-Instruct, MemSuit reaches 35.56 average F1 and 27.73 BLEU, beating SimpleMem’s 30.09 F1 and 22.24 BLEU.

Benchmark

GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations

Jingbo Yang, Kwei-Herng Lai et al.

arXiv 2026 · 2026

GroupMemBench combines Graph-grounded Message Synthesis, Path-Based Context Sampling, Message Generation, Adversarial Question Synthesis, and a Solve–Judge–Refine Loop to stress-test group memory in LLM agents. On GroupMemBench, Hindsight reaches 46.01% average accuracy while a simple BM25 baseline attains 43.22%, revealing that current memory ingestion pipelines fail to preserve crucial multi-user structure.

Benchmark

H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions

Shiping Zhu, Yibo Yang et al.

arXiv 2026 · 2026

H2HMem constructs dyadic and multi-party multimodal conversations via a Participant Profile Generation, Scenario Construction, Image Collection and Human Refinement, Image Captioning and Dialogue Generation, and Question-Answer Pairs Construction pipeline, then evaluates agents on recall, reasoning, and application tasks. On the H2HMem benchmark, the best configuration with A-Mem on GPT-4.1-Nano reaches an overall LLM-as-Judge score of 0.5757, revealing large gaps versus ideal memory behavior.

Agent Memory

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

Wei-Chieh Huang, Weizhi Zhang et al.

arXiv 2026 · 2026

Harness the Memory instruments a Unified Evaluation Harness, External Memory Substrate Families, Internal Memory Substrate Families, Retrieval Depth Sweep, and Scalability Study to compare 11 memory substrates under identical agents. Harness the Memory shows, for example, that M7 Distilled Strategies reaches 32.1% TSR on ALFWorld-unseen with QWEN3-32B-AWQ, a +9.7 percentage point gain over the NoMem baseline at only 1.23× latency.

Agent Memory

HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory

Yuanhua Lin, Yile Li et al.

arXiv 2026 · 2026

HERO organizes dialogue into Episodic Traces, Episodic Units, Episodic Cues, Profile Insights, and Profile Cues inside a heterogeneous memory graph, then runs profile-aware cue activation and Personalized PageRank for retrieval. On LoCoMo, HERO attains 56.06% F1 and 87.99% accuracy, beating EverMemOS by +10.58 F1 and +3.96 accuracy.

RAGLong-Term Memory

HingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable Dialogues

Yijie Zhong, Yunfan Gao, Haofen Wang

· 2026

HingeMem combines Boundary Guided Long-Term Memory, Dialogue Boundary Extraction, Memory Construction, Query Adaptive Retrieval, Hyperedge Rerank, and Adaptive Stop to segment dialogues into element-indexed hyperedges and plan query-specific retrieval. On LOCOMO, HingeMem achieves 63.9 overall F1 and 75.1 LLM-as-a-Judge score, surpassing the best baseline Zep (56.9 F1) by 7.0 F1 without using category-specific QA formats.

Benchmark

HippoCamp: Benchmarking Contextual Agents on Personal Computers

Zhe Yang, Shulin Tian et al.

arXiv 2026 · 2026

HippoCamp evaluates contextual agents on realistic personal file systems using structured trajectories with search, perception, and reasoning capability tags over 42.4 GB of multimodal data. On HippoCamp, ChatGPT Agent Mode reaches 48.3% profiling accuracy and 62.8% factual retention accuracy, while RAG and search agents lag far behind.

Cognitive ArchitectureLong-Term Memory

Human-Like Lifelong Memory: A Neuroscience-Grounded Architecture for Infinite Interaction

Diego C. Lerma-Torres

· 2026

Human-Like Lifelong Memory combines Executive Function and Working Memory, a Memory Service Knowledge Graph, and a Thalamic Gateway to implement dual-process, valence-aware lifelong memory. Human-Like Lifelong Memory is a theoretical framework with seven functional properties and testable predictions rather than benchmark numbers against specific baselines.