Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 11 of 19

RAG

RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation

Kyle Wild, Yusuke Takahashi, Asako Uraki

arXiv 2026 · 2026

Ingest-Time Semantic Compilation (ISC) builds a semantic substrate in PostgreSQL using a geometric layer, symbolic layer, validation gate, and index_outbox for incremental maintenance and migration. On 500 broadcast-interview transcripts, ISC’s compiled claims reach 85.2% accuracy from roughly 2.2k reader tokens, beating the best chunk configuration at 72.5% from roughly 16.3k tokens and matching a contextualized stack that spends ~47.7k tokens.

Long-Term Memory

REAL: A Reasoning-Enhanced Graph Framework for Long-Term Memory Management of LLMs

Keer Lu, Liwei Chen et al.

arXiv 2026 · 2026

REAL represents conversational memory as a temporal directed property graph using Atomic Fact Extraction, Incremental Graph Update with Non-Destructive Temporal Evolution, Confidence Stratification and Upgrade, and Exploration Intent Enrichment. On LoCoMo, REAL with DeepSeek-V3 reaches 59.98% EM versus 55.12% EM for A-MEM, and averages a 22.72% improvement over existing memory management methods.

Agent Memory

RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction

Haonan Bian, Zhiyuan Yao et al.

arXiv 2026 · 2026

RealMem constructs realistic long-term project dialogues via Project Foundation Construction, Multi-Agent Dialogue Generation, and Memory and Schedule Management over eleven scenarios and 2,000+ cross-session dialogues. On RealMem, the Oracle QA Score reaches 0.804 while the strongest memory system, MemoryOS, achieves only 0.567, quantifying the difficulty of real-world project memory.

Agent Memory

RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

Mihir Shriniwas Arya

arXiv 2026 · 2026

RECON constructs long-context cases via attribute-conditioned blueprint synthesis, a provenance DAG and proof-trace grounding, and deterministic task synthesis over evolving evidence. On 1,414 clean questions, RECON shows an Oracle at 54.6% Accuracy while the best non-Oracle system (Gemini-2.5-Pro) reaches 22.4% Accuracy, exposing large retrieval and reasoning gaps.

Benchmark

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Zhaochen Yu, Yingcheng Wu et al.

arXiv 2026 · 2026

Recuris maintains a Working Memory state, invokes Experiential Memory skills via an invocation policy 𝜌𝑘, and verifies progress with checker set C𝑘 inside a recursive Skill Memory M𝑘. On 𝜏2-Retail, Recuris lifts doubao-2.0-pro from 58.1% to 81.4% task success (+23.3 points) and GPT-5.6 Sol from 58.3% to 76.1% (+17.8 points).

Long-Term Memory

Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents

Wanqi Zhou, Jiawei Lu et al.

arXiv 2026 · 2026

RIME combines Question-guided evidence recall, Historical memory grounding, Joint memory formation and reconciliation, and Selective source-context recovery into a retrieval-induced long-term memory bank for LLM agents. On the 1,540-question LoCoMo benchmark, RIME achieves a Judge score of 84.29 with GPT-5.6 Sol, surpassing Nemori’s 83.05 while cutting query-time tokens from 17.74K to 5.21K.

Long-Term Memory

Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents

Baichuan Li, Junyi Yao, Zihao Zheng

arXiv 2026 · 2026

Memory–Clarification Boundary (MCB) uses persist, ephemeral, verify, and clarify actions plus the MCB-Act tool-call variant to probe memory commitment decisions in LLM agents. On the 70-item held-out MCB test, Qwen3.5-9B few-shot prompting raises accuracy from 0.557 to 0.771 while Qwen3.5-9B act-mode accuracy falls to 0.343, showing that stated decisions and tool-call choices diverge.

Long-Term Memory

REMem: Reasoning with Episodic Memory in Language Agent

Yiheng Shu, Saisri Padmaja Jonnalagedda et al.

· 2026

REMem converts interaction histories into a hybrid memory graph via Gist Extraction, Fact Extraction, and Graph Construction, then queries it with Agentic Inference tools like semantic retrieve and find entity contexts. On the Test of Time benchmark, REMem-I achieves 93.1% EM compared to 66.9% for HippoRAG 2, a +26.2 point gain in episodic reasoning.

Agent Memory

Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey

Wei-Chieh Huang, Weizhi Zhang et al.

arXiv 2026 · 2026

Rethinking Memory Mechanisms of Foundation Agents in the Second Half organizes foundation agent memory using a three-dimensional taxonomy of Memory Substrates, Memory Cognitive Mechanisms, and Memory Subjects plus an operation and optimization view. Rethinking Memory Mechanisms of Foundation Agents in the Second Half synthesizes 218 memory-related agent papers from 2023 Q1–2025 Q4, highlighting the sharp acceleration of external memory and working or episodic mechanisms in 2025.

Long-Term Memory

Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems

Yi Ting Shen, Kentaroh Toyoda, Alex Leung

arXiv 2026 · 2026

Revocation Enforcement Study instruments five agent-memory systems and a retrieval-time guard around their default retrieval and revocation mechanisms. It shows that Graphiti and mem0 (expiry override) still return revoked policies ranked first, leading agents to unsafe actions in 44.2% and 42.1% of trials, while a store-level filter or guard eliminates these failures.

Benchmark

RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory

Jingbo Ji, Lingyi Li et al.

arXiv 2026 · 2026

RippleMem converts dialogues into Cue-Rich Episodic Memory Construction, organizes them in an Event-Centric Memory Graph, and queries them via Adaptive Associative Recollection and Evidence Assembly. On LoCoMo, RippleMem attains 87.14% LLM-as-a-Judge accuracy, 52.49% F1, and 44.05 BLEU-1, beating RF-Mem’s 83.83% judge accuracy.

Agent Memory

RoutePrism: Tracing Construction Order Effects in Agent Memory

Dong Xu, Zhangfan Yang et al.

arXiv 2026 · 2026

RoutePrism traces memory construction using Task-preserving eligibility, Paired Memory Construction, Observable Diagnostics, and Support Intervention and Analysis to compare forward and reversed processing orders over the same record pool. On PersonaMem-32K, RoutePrism’s support restoration recovers over 60 percentage points of lost accuracy under both focal compaction and bounded recency compared to survivor-selection baselines.

Long-Term Memory

RUMBA: Russian User Memory Benchmark

Elizaveta Shevtsova, Inna Glebkina et al.

arXiv 2026 · 2026

RUMBA organizes long-term conversational memory evaluation using a three-axis taxonomy, timestamped user–assistant dialogues, and a two-stage add plus eval pipeline. On the RUMBA benchmark, full-context gpt-4.1-mini scores 60.08% vs an overall Agent/RAG mean of 55.80% in Russian, revealing distinct failure modes across temporal and multi-session slices.

Memory Architecture

Safin-1: Safety from Within through Memory-Native State Evolution

Ming Zhang, Kaisen Yang et al.

arXiv 2026 · 2026

Safin-1 routes Context-Derived State Anchors, Content-Routed State Retrieval, Persistent Capability States, and an Efficient Producer–Reader Implementation through MARCH to make safety a memory-native capability. Safin-1 reaches 80.1 AIME 2025 Avg@64 vs 71.9 for Qwen3.5 and reduces Average ASR from 7.25% to 3.84% using its Safety State.

Benchmark

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Lingyang Zeng, Guangze Chen et al.

arXiv 2026 · 2026

Setoka builds a UUQA pipeline with User Understanding Hierarchy, correlation-aware personality trait sampling, psychological-scale-based behavior pattern generation, and an event-grounded generation tree over heterogeneous data. On Setoka, the best configuration reaches 0.85 on Semantic Memory but only 0.24 on Personality Traits, revealing large gaps in hierarchical user understanding.

Benchmark

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

Yingtie Lei, Zhongwei Wan et al.

arXiv 2026 · 2026

SkillEvolBench orchestrates Environment-specific Shared Skill Library, Skill Author, Episodic Task Attempts, and Structured Verifier Feedback to turn trajectories into external procedural skills and then freeze them. SkillEvolBench’s main result is that Raw-Trajectory achieves mean ESR 37.6% and ARSR 44.7%, while most curated and self-generated skill variants fail to match these deployment scores.

Agent Memory

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

Huacan Chai, Yukai Wang et al.

arXiv 2026 · 2026

SMMBench evaluates multimodal agent memory across independently originated sources using components like QA Preparation, Conversational Source Synthesis, Source-Aware Evidence Insertion, Memory Construction, and Evaluation Framework. On 1877 samples from 264 sources, SMMBench’s GOLDEN EVIDENCE BASELINE reaches 0.7473 overall, while the best baseline HMRAG only achieves 0.4933.

Benchmark

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

Fengrong Wan, Chengcan Wu, Ningtao Lyu

arXiv 2026 · 2026

SodaMem ingests dialogue into FactEvents, stores them in a temporal graph with hybrid BM25–dense indexes, and answers via a planner–reader loop over multi-tunnel retrieval. On LongMemEval-S, SodaMem achieves 92.8% accuracy (464/500) with deepseek-v4-flash, beating GPT-4o-mini pipelines like MemOS (77.8%) by +15.0 percentage points at comparable or lower cost.

Benchmark

STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?

Hanxiang Chao, Yihan Bai et al.

arXiv 2026 · 2026

STALE evaluates latent user-state tracking by probing State Resolution, Premise Resistance, and Implicit Policy Adaptation across 400 conflict scenarios packaged into long user-assistant histories. CUPMEM applies structured state consolidation and propagation-aware search on STALE, reaching 68.0% overall accuracy versus 55.2% for Gemini-3.1-pro.