Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 9 of 19

RAGBenchmarkLong-Term Memory

MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents

Shu Wang, Edwin Yu et al.

· 2026

MemMachine combines Short-term memory, Long-term memory, Profile memory, and the Retrieval Agent to store raw conversational episodes and retrieve clustered context around nucleus matches. On LoCoMo, MemMachine scores 0.9169 with gpt-4.1-mini while using about 80% fewer input tokens than Mem0, and reaches 93.0% on LongMemEvalS with GPT-5-mini.

Benchmark

MEMO: Multimodal Evidence Memory Organization for Long-Horizon LLM Agents

Xian Gao, Jinpeng Wang et al.

arXiv 2026 · 2026

MEMO builds query-specific working memory by combining a query-conditioned evidence extractor, unit-materialization, a memory manager, and deterministic working memory building into a multimodal text plus visual pipeline. Under a 128-token budget on 2WikiMultiHopQA with Qwen3-VL-32B, MEMO reaches 73.91 F1, a +17.65 gain over Text-only memory at 56.26 F1.

Long-Term Memory

MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations

Xixuan Hao, Zeyu Zhang et al.

arXiv 2026 · 2026

MemOps decomposes long-horizon memory into structured Remember, Forget, Update, Reflect, and TrajectoryOps traces with six probe types like OperationTrace and StateTransition. On the MemOps benchmark, Claude-Sonnet-4.5 attains 0.916 adjacent accuracy versus Temp-LoRA’s 0.162, exposing deep reliability gaps across memory paradigms.

RAGAgent MemoryLong-Term MemoryMemory Architecture

Memory as Metabolism: A Design for Companion Knowledge Systems

Stefan Miteski

· 2026

Memory as Metabolism defines companion knowledge systems with five retention operations (TRIAGE, DECAY, CONTEXTUALIZE, CONSOLIDATE, AUDIT) plus memory gravity and minority-hypothesis retention over a raw buffer, active wiki, and cold memory. Instead of benchmark gains, Memory as Metabolism’s main result is a governance specification that separates descriptive, taxonomic, and normative claims and predicts improved coherence stability, fragility resistance, monoculture resistance, and effective minority-hypothesis influence for companion wikis.

Memory Architecture

Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

Simeng Zhang, Yilong Chen et al.

arXiv 2026 · 2026

Memory-Augmented Compression builds a memory bank, uses a memory retriever, performs memory-augmented prefill, and runs memory-guided compressed inference to support short Chain-of-Thought reasoning. On GSM8K, Memory-Augmented Compression with CoD achieves 89.3% accuracy versus 67.9% for CoD alone, while keeping 1.49× lower latency than standard CoT.

BenchmarkBenchmarkAgent Memory

MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization

Weizhi Zhang, Xiaokai Wei et al.

· 2026

MEMORYCD builds a user memory pool Mu from lifelong Amazon Review histories and evaluates long-context prompting, Mem0, LoCoMo, ReadAgent, MemoryBank, and A-Mem across rating, ranking, and personalized text tasks. On Books and Home & Kitchen, MEMORYCD shows GPT-5 reaches RMSE 0.551–0.624 and NDCG@3 up to 0.610, while Gemini-2.5 Pro peaks at ROUGE-L 0.222 for generation, revealing substantial remaining gaps to real user behavior.

SurveyRAGAgent Memory

Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers

Pengfei Du

· 2026

Memory for Autonomous LLM Agents decomposes agent memory into a POMDP-grounded write–manage–read loop, a three-dimensional taxonomy, and five mechanism families spanning context compression, retrieval stores, reflection, hierarchical virtual context, and policy-learned management. Memory for Autonomous LLM Agents synthesizes results like Voyager’s 15.3× tech-tree speedup and MemoryArena’s 80%→45% drop to show that memory architecture often matters more than backbone choice.

Benchmark

Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning

Zhen Bi, Xueshu Chen et al.

arXiv 2026 · 2026

Memory Boundary-Aware Router combines Behavioral Scientific Knowledge Boundary Characterization, Knowledge-Circuit View of Internal Knowledge Boundaries, External Boundary-Aware Data Router, and Internal Boundary-Aware Parameter Router to control conditional memory inside Transformer blocks. On BioProBench Error Correction, Memory Boundary-Aware Router reaches 0.63 accuracy with Qwen3-8B, surpassing the 0.62 accuracy of Qwen3-8B-Memory-LoRA.

Agent Memory

MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends

Chaoqun Zhan, Qiang Zhou et al.

arXiv 2026 · 2026

MemoryLake organizes confirmed conclusions, supporting evidence, and reusable experience into separate tracks with presence policies, using gpt-5-mini and bge-m3 for consolidation, retrieval, and bounded prompt assembly. On the shared MemoryArena sets, MemoryLake reaches a 20.5% equal-weight macro-average SR compared to the 13.6% best comparator, with SR 9/40 in mathematics, 12/20 in physics, and 4/20 in progressive retrieval.

Long-Term Memory

Memory Poisoning Attack and Defense on Memory Based LLM-Agents

Balachandra Devarangadi Sunil, Isheeta Sinha et al.

· 2026

Memory Poisoning Attack and Defense on Memory Based LLM-Agents combines Input Output Moderation, Memory Sanitization with trust-aware retrieval, bridging steps, and indication prompts to study and harden long-term memory in EHR agents. On MIMIC-III with GPT-4o-mini and Llama-3.1-8B-Instruct, Memory Poisoning Attack and Defense on Memory Based LLM-Agents shows that adding realistic initial memory drops GPT-4o-mini ASR from 62% to 6.67% and ISR from 100% to 26.67% compared to the empty-memory baseline.

Long-Term Memory

MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models

Zecheng Tang, Baibei Ji et al.

· 2026

MemoryRewardBench constructs paired memory management trajectories across Sequential Pattern, Parallelism Pattern, Mixed Pattern, long-context reasoning, multi-turn dialogue, and long-form generation to test 13 reward models. MemoryRewardBench shows Claude-Opus-4.5 at 74.75 average accuracy and GLM4.5-106A12B at 68.21 on its 2,400-example benchmark, exposing generational gains over larger predecessors.

Agent Memory

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery

Enze Ma, Yufan Zhou et al.

arXiv 2026 · 2026

MEMPROBE combines Simulated Users and Leak-Controlled Tasks, Simulation Rollout, and Recovery Scoring to audit what user state memory agents actually preserve. On 1,550 hidden user-state targets, MEMPROBE finds category-balanced reconstruction only around 0.624 for longctx_full under full-store access and as low as 0.473 for mem0 under top-k retrieval, even though all systems reach ~99.9% task completion.

Agent Memory

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

Chunyu Li, Jingyi Kang et al.

arXiv 2026 · 2026

MemReranker connects Qwen3-Reranker with multi-teacher label generation, BCE pointwise distillation, InfoNCE contrastive fine-tuning, and multi-turn dialogue data engineering to build reasoning-aware memory rerankers. On LongMemEval, MemReranker-4B attains MAP 0.8043 versus 0.7259 for Gemini-3-Flash, while MemReranker-0.6B matches or exceeds GPT-4o-mini at ∼8× lower latency.

Memory Architecture

MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents

Tianyu Hu, Weikai Lin et al.

· 2026

MemRouter combines a Memory Router Architecture, Memory Store and Retrieval, and Answer Agent built on a frozen Qwen2.5-7B backbone to make turn-level memory admission decisions in embedding space. On LoCoMo, MemRouter achieves 52.0 overall F1 with Qwen2.5-7B-Instruct, beating Memory-R1-GRPO at 43.1 F1 under a like-for-like backbone and reaching 55.5 F1 when swapping only the answer agent to Qwen3.5-35B-A3B.

Agent Memory

MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair

Xuanze Chen, Xukang Xie et al.

arXiv 2026 · 2026

MemSecBench links Lifecycle Task Packages, Build-MemSecBench-Case Skill, Lifecycle Evaluation Workflow, Evidence-Based LLM Judging, and a 24-configuration matrix to trace malicious semantics through memory systems. MemSecBench reports 50.3% End-to-End Attack Success Rate and 56.1% Selective Repair Success Rate across 310 cases and 24 configurations, contrasting Native with Mem0, Mem0-Graph, and A-MEM.

Long-Term Memory

MemSentry: A Framework for Detecting Persistent Memory Poisoning in Agentic AI

Ayan Roy, Kaustuvi Basu

arXiv 2026 · 2026

MemSentry evaluates persistent memory operations via a 9-step pipeline using components like Source Trust, Semantic Classification, Attack Radius, and Access Risk over a dependency DAG. On a 1,000-scenario GPT-4 dataset, MemSentry with SBERT+LR achieves 91.7% accuracy and 0.908 macro-F1, detecting 100% external quarantine-class threats compared to Regex, TF-IDF+SVM, and SetFit.

Agent Memory

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Zhishang Xiang, Zerui Chen et al.

arXiv 2026 · 2026

MemSyco-Bench evaluates agent memory use across Objective Fact Judgment, Contextual Scope Control, Memory-Evidence Conflict, Personalized Memory Use, and Valid Memory Selection, using a structured memory-decision schema, question instantiation, long-term dialogue simulation, and multi-stage quality validation pipeline. On MemSyco-Bench, existing memory systems like SuperMemory and MemGPT frequently degrade accuracy (e.g., -23.12 points) and increase sycophancy rates (e.g., +37.24 points) compared to no-memory baselines.

Benchmark

MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

Arseniy Varlamov, Rishat Zinnatullin et al.

arXiv 2026 · 2026

MemToC builds a controlled arbitration pipeline from ToolHop pool construction and filtering, Distractor generation, Correctness-conditioned evaluation, and Annotator quality control to separate source correctness from source preference. On MemToC, cross-fitted SFT lifts Llama-3.1-8B-Instruct correct-answer retention from 17.1% to 31.6% while keeping correct-tool following at 92.8–93.1%, and DPO achieves 22.3% retention with 92.1% correct-tool following.