Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 4 of 19

Long-Term Memory

Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory

Sahil Sen, Elias Lumer et al.

· 2026

Chronos decomposes dialogue into structured events via the Event Extraction pipeline, stores them in dual event calendar and turn calendar indexes, and uses Dynamic Prompting, Initial Retrieval, and the Chronos Agent for temporal-aware tool-calling. On LongMemEvalS, Chronos Low reaches 92.60% overall accuracy and Chronos High 95.60%, beating EmergenceMem Internal by 7.67 percentage points and Mastra’s OM by 3.02 points.

BenchmarkAgent Memory

ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents

Mofasshara Rafique, Laurent Bindschaedler

· 2026

ClawVM manages agent state as typed pages via the SessionPageTable, RepresentationSelector, FaultObserver, WritebackJournal, and ClawVMEngine inside the agent harness. Across four OpenClaw-derived workloads and six token budgets, ClawVM cuts explicit faults from 67.8 (retrieval baseline) and 1.5 (Compaction-Hybrid) to 0.0 while adding median <50 μs policy-engine overhead per turn.

Long-Term Memory

CloneMem: Benchmarking Long-Term Memory for AI Clones

Sen Hu, Zhiyu Zhang et al.

· 2026

CLONEMEM builds AI clone memory from Persona and Macro-Level Life Arcs, Meso-Level Phase Generation, Micro-Level Digital Trace Generation, and Evaluation Question Construction over synthetic diaries, posts, and emails. On CLONEMEM, flat retrieval achieves Recall-Any-Any 0.6103 at k=20 versus 0.3913 for A-Mem, revealing a validity–fidelity trade-off in existing memory systems.

Long-Term Memory

CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning

Yubo Wang, Qiuyu Zhao et al.

arXiv 2026 · 2026

CMI-Mem manages long-term dialogue memory via a Structured Multi-dimensional Memory Architecture, Action-conditioned CMI Reward, and Reinforcement Training with GRPO that jointly optimize what to store, update, or discard. On LoCoMo, CMI-Mem-4B achieves 72.1% accuracy compared to 68.1% for MemBuilder-RL-4B, and on MemoryAgentBench CMI-Mem-8B reaches 56.7 vs 51.7 for MemBuilder-RL-8B.

Benchmark

Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP

Martin Vogel, Falk Meyer-Eschenbach et al.

· 2026

Codebase-Memory parses repositories with a multi-pass pipeline using the Parse stage, Build stage, Serve stage, FunctionRegistry, Louvain communities, and MCP tool interface to build a persistent SQLite knowledge graph. On a 31-language benchmark, Codebase-Memory reaches 0.83 quality versus 0.92 for an Explorer Agent while using ten times fewer tokens and 2.1 times fewer tool calls.

Agent Memory

CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems

Chengxin Yu, Zhaoxin Fan et al.

arXiv 2026 · 2026

CoMem combines Private Experience Sedimentation, Collective Wisdom Curation, Parallel Dual-Stream Retrieval, and a strict empirical promotion gate into a two-tier group memory for multi-agent systems. On ALFWorld with MacNet, CoMem reaches 89.55% success versus the 79.85% No-memory baseline, and also lifts PDDL success from 60.78% to 70.19%.

Benchmark

Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

Jiahe Geng, Jinpeng Wang, Kun Yuan

arXiv 2026 · 2026

RSM-full combines Cosine-gated max-member merge, Atom-aware grouped context assembly, |v⊤₁,m q| retrieval, and higher-rank storage variants into an online clustered-memory pipeline for LLM agents. On AMA-Bench, RSM-full achieves 0.311 average correctness at ~4k tokens, beating Online K-Means by up to +6.0 pp and Budget-RAG by +1.9 pp.

Benchmark

Consolidator: Learning Persistent Routed Memory Across Context Boundaries

Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung

arXiv 2026 · 2026

Consolidator attaches a shared slot-local Consolidator, hierarchical router, routed short-term memory, persistent long-term memory, and sliding-attention KV cache to PMNet to learn replay-free STM-to-LTM consolidation. On the two-segment modulo-10 same-address update task, Consolidator reaches 87.02% updated-mapping LTM recall versus 18.32% for forced identity accumulation using only 12.35K trainable parameters.

Long-Term Memory

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

Zhuoshi Pan, Qizhi Pei et al.

arXiv 2026 · 2026

ContextPilot augments Perception & Planning, Information Retrieval, Memory Management, and Context Offloading tools, then trains them with context-aware partial rollout and fine-grained credit assignment. On long-context QA, ContextPilot-14B-RL scores 72.20 average across NovelQA, ∞Bench, LongMemEval-S, and BrowseComp+ compared to 70.11 for StateLM-14B-RL.

Benchmark

ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents

Yating Wu, Yuhao Zhang et al.

· 2026

ContextWeaver organizes agent histories using Dependency-Aware Context Construction, Dependency Summarizer, and a Validation and Test Layer into a dependency graph that drives selective context weaving. On SWE-Bench Verified, ContextWeaver with Claude Sonnet 4 achieves 66.0% pass@1 versus 63.2% for the Sliding Window baseline under the same window size.

RAG

ConvMemory v2: A Recall-Preserving Top-10 Evidence Reranker for Conversational Memory Retrieval

Taiheng Pan

arXiv 2026 · 2026

ConvMemory v2 adds a recall-preserving cascade on top of ConvMemory v1, using a fine-tuned ms-marco-MiniLM-L-6-v2 cross-encoder, a strict anti-shortcut inference contract, and a token-evidence ablation to reorder only v1’s protected top-10 candidates. On the LoCoMo benchmark, ConvMemory v2 raises FULL MRR from 0.5824 to 0.6560 over ConvMemory v1 and comes within 0.013 MRR of mxbai-rerank-large-v1 while running about 68× cheaper than the full-pool cross-encoder.

Benchmark

CoreMem: Riemannian Retrieval and Fisher-Guided Distillation for Long-Term Memory in Dialogue Agents

Jiaqi Chen, Yongqin Zeng et al.

arXiv 2026 · 2026

CoreMem combines Riemannian retrieval, Fisher-guided discrete token distillation, Residual Metric Fusion, and an Edge-Cloud Hybrid Architecture to retrieve and compress long-term dialogue memories under strict VRAM budgets. On LOCOMO with MiniLM-L6, CoreMem-Fusion reaches Judge accuracy 0.540 versus 0.531 for NaiveRAG, with +4.51 pp Open-domain and +4.17 pp Temporal gains.

Agent Memory

COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents

Hongji Pu, Ruixiang Tang, Yongfeng Zhang

arXiv 2026 · 2026

COUNTERMEM combines Module I: Propose Alternatives, Module II: Check Alternatives and Store Corrections, and Module III: Retrieve and Select Records to turn failed actions into reusable, verified counterfactual memory. On the four-domain suite with gpt-oss-120b and ReAct, COUNTERMEM reaches a 57.5 average score versus 39.0 for ReAct without augmentation (+18.5 points) and reduces task-run tokens from 6.0M to 4.4M.

Long-Term Memory

CreaMem: A Scene-Aware Memory Architecture for Personalized Agents

Qixuan Sun, Yue Que et al.

arXiv 2026 · 2026

CreaMem organizes long-term agent memory into Meta Memory Manager, Episodic Memory, three Life Scene Memories, and a Core Memory, with a Planner LLM orchestrating per-memory balanced retrieval. On LongMemEval-S, CreaMem achieves 66.40% 4o-Judge accuracy versus 60.20% for MemGAS, and on LoCoMo reaches 54.61% versus 45.62% for HippoRAG 2.

Memory Architecture

CrystalMem: Elastic Memory for Self-Evolving LLM Agents via Knowledge Crystallization

Beining Wu, Jun Huang

arXiv 2026 · 2026

CrystalMem combines Monitor, Crystallize, and Recrystallize around a four-state fidelity ladder, driven by advantage-weighted influence with dependency coupling and guarded by a verification gate. CrystalMem reaches 57.4 average restored capability across seven environments, beating R3Mem at 52.8 (+4.6 pp) under the same elastic budget.

Long-Term Memory

CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations

Yulin Hu, Yanyan Zhao et al.

arXiv 2026 · 2026

CUE-Mem organizes synthetic multimodal user histories through Script Synthesis, Multimodal Dialog Synthesis, Data Review, and four tasks: Entity Recall, Long Pattern, Personalized Recommendation, and Answer Refusal. On CUE-Mem, textualized memory systems reach implicit Memory Avg. only 38.5 vs 80.8 with Oracle Evidence, exposing a large preservation and retrieval bottleneck for implicit cues.

Cognitive ArchitectureAgent Memory

D-Mem: A Dual-Process Memory System for LLM Agents

Zhixing You, Jiachen Yuan, Jason Cai

· 2026

D-Mem combines Mem0∗, Quality Gating, and Full Deliberation into a dual-process memory system that incrementally stores vector memories and selectively scans raw history. On LoCoMo with GPT-4o-mini, D-Mem’s Quality Gating reaches 53.5 F1 versus the Mem0∗ baseline’s 51.2 F1, recovering 96.7% of the 55.3 F1 Full Deliberation performance with far fewer tokens.

Benchmark

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Ankit Goyal, Jaideep Ray

arXiv 2026 · 2026

Does Your Agent’s Memory Survive a Model Upgrade? decomposes agent memory into LC-RAW, RAG, NOTES, and KG-fixed with separate writer, reader, and embedder roles. On 48 synthetic histories, KG-fixed changes by only +0.0004 ± 0.0020 accuracy after a writer swap, while NOTES can shift by +9.91 or −13.28 percentage points and a 50/50 mixed embedding index forfeits most of an 11.90-point re-embedding gain.

Survey

DolphinBench: Mapping the Pareto Frontier of Agent Memory

Soumil Rathi, Deshraj Yadav, Taranjeet Singh

· 2026

DolphinBench builds Personas, History Planning, History Generation, Test Construction, and Test Verification into a hierarchical pipeline to create long-horizon, action-based memory tasks. On DolphinBench’s 600 tasks, the Hermes + GPT-5.6-Luna + Mem0 configuration reaches 70.67% accuracy while exposing distinct cost and latency tradeoffs against built-in memory and other baselines.

Long-Term Memory

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

Jong Wook Kim, Byoungjae Min et al.

arXiv 2026 · 2026

DP-MemView uses Slot-Level View Selection, Attribute Accounting, and the DP-MemView Interface Contract to turn raw memories into differentially private public views for response LLMs. On the PairedMem benchmark with Qwen2.5-7B, DP-MemView (on) achieves AUC 0.490 vs 0.842 for RawReadSet while maintaining U 0.877.