Paper archive

All AI memory papers

Browse 314 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 5 of 14

Agent Memory

MemLineage: Lineage-Guided Enforcement for LLM Agent Memory

Ciyan Ouyang, Rui Hou

arXiv 2026 · 2026

MemLineage wraps a single memory store with six modules: Provenance metadata, Ed25519 signing, an RFC 6962 Merkle log, a weighted lineage DAG, verifier-aware retrieval, and a sensitive-action gate that enforces Untrusted-Path Persistence. On a deterministic harness with AgentPoison-style, MemoryGraft-style, and sleeper-via-derivation attacks, MemLineage is the only configuration that achieves 0.00 ASR on all three while keeping per-operation overhead below one millisecond.

RAGBenchmarkLong-Term Memory

MemMachine: A Ground-Truth-Preserving Memory System for Personalized AI Agents

Shu Wang, Edwin Yu et al.

· 2026

MemMachine combines Short-term memory, Long-term memory, Profile memory, and the Retrieval Agent to store raw conversational episodes and retrieve clustered context around nucleus matches. On LoCoMo, MemMachine scores 0.9169 with gpt-4.1-mini while using about 80% fewer input tokens than Mem0, and reaches 93.0% on LongMemEvalS with GPT-5-mini.

RAGAgent MemoryLong-Term MemoryMemory Architecture

Memory as Metabolism: A Design for Companion Knowledge Systems

Stefan Miteski

· 2026

Memory as Metabolism defines companion knowledge systems with five retention operations (TRIAGE, DECAY, CONTEXTUALIZE, CONSOLIDATE, AUDIT) plus memory gravity and minority-hypothesis retention over a raw buffer, active wiki, and cold memory. Instead of benchmark gains, Memory as Metabolism’s main result is a governance specification that separates descriptive, taxonomic, and normative claims and predicts improved coherence stability, fragility resistance, monoculture resistance, and effective minority-hypothesis influence for companion wikis.

Memory Architecture

Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning

Simeng Zhang, Yilong Chen et al.

arXiv 2026 · 2026

Memory-Augmented Compression builds a memory bank, uses a memory retriever, performs memory-augmented prefill, and runs memory-guided compressed inference to support short Chain-of-Thought reasoning. On GSM8K, Memory-Augmented Compression with CoD achieves 89.3% accuracy versus 67.9% for CoD alone, while keeping 1.49× lower latency than standard CoT.

BenchmarkBenchmarkAgent Memory

MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization

Weizhi Zhang, Xiaokai Wei et al.

· 2026

MEMORYCD builds a user memory pool Mu from lifelong Amazon Review histories and evaluates long-context prompting, Mem0, LoCoMo, ReadAgent, MemoryBank, and A-Mem across rating, ranking, and personalized text tasks. On Books and Home & Kitchen, MEMORYCD shows GPT-5 reaches RMSE 0.551–0.624 and NDCG@3 up to 0.610, while Gemini-2.5 Pro peaks at ROUGE-L 0.222 for generation, revealing substantial remaining gaps to real user behavior.

SurveyRAGAgent Memory

Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers

Pengfei Du

· 2026

Memory for Autonomous LLM Agents decomposes agent memory into a POMDP-grounded write–manage–read loop, a three-dimensional taxonomy, and five mechanism families spanning context compression, retrieval stores, reflection, hierarchical virtual context, and policy-learned management. Memory for Autonomous LLM Agents synthesizes results like Voyager’s 15.3× tech-tree speedup and MemoryArena’s 80%→45% drop to show that memory architecture often matters more than backbone choice.

Agent Memory

MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends

Chaoqun Zhan, Qiang Zhou et al.

arXiv 2026 · 2026

MemoryLake organizes confirmed conclusions, supporting evidence, and reusable experience into separate tracks with presence policies, using gpt-5-mini and bge-m3 for consolidation, retrieval, and bounded prompt assembly. On the shared MemoryArena sets, MemoryLake reaches a 20.5% equal-weight macro-average SR compared to the 13.6% best comparator, with SR 9/40 in mathematics, 12/20 in physics, and 4/20 in progressive retrieval.

Long-Term Memory

Memory Poisoning Attack and Defense on Memory Based LLM-Agents

Balachandra Devarangadi Sunil, Isheeta Sinha et al.

· 2026

Memory Poisoning Attack and Defense on Memory Based LLM-Agents combines Input Output Moderation, Memory Sanitization with trust-aware retrieval, bridging steps, and indication prompts to study and harden long-term memory in EHR agents. On MIMIC-III with GPT-4o-mini and Llama-3.1-8B-Instruct, Memory Poisoning Attack and Defense on Memory Based LLM-Agents shows that adding realistic initial memory drops GPT-4o-mini ASR from 62% to 6.67% and ISR from 100% to 26.67% compared to the empty-memory baseline.

Long-Term Memory

MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models

Zecheng Tang, Baibei Ji et al.

· 2026

MemoryRewardBench constructs paired memory management trajectories across Sequential Pattern, Parallelism Pattern, Mixed Pattern, long-context reasoning, multi-turn dialogue, and long-form generation to test 13 reward models. MemoryRewardBench shows Claude-Opus-4.5 at 74.75 average accuracy and GLM4.5-106A12B at 68.21 on its 2,400-example benchmark, exposing generational gains over larger predecessors.

Agent Memory

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

Chunyu Li, Jingyi Kang et al.

arXiv 2026 · 2026

MemReranker connects Qwen3-Reranker with multi-teacher label generation, BCE pointwise distillation, InfoNCE contrastive fine-tuning, and multi-turn dialogue data engineering to build reasoning-aware memory rerankers. On LongMemEval, MemReranker-4B attains MAP 0.8043 versus 0.7259 for Gemini-3-Flash, while MemReranker-0.6B matches or exceeds GPT-4o-mini at ∼8× lower latency.

Memory Architecture

MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents

Tianyu Hu, Weikai Lin et al.

· 2026

MemRouter combines a Memory Router Architecture, Memory Store and Retrieval, and Answer Agent built on a frozen Qwen2.5-7B backbone to make turn-level memory admission decisions in embedding space. On LoCoMo, MemRouter achieves 52.0 overall F1 with Qwen2.5-7B-Instruct, beating Memory-R1-GRPO at 43.1 F1 under a like-for-like backbone and reaching 55.5 F1 when swapping only the answer agent to Qwen3.5-35B-A3B.

Agent Memory

MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair

Xuanze Chen, Xukang Xie et al.

arXiv 2026 · 2026

MemSecBench links Lifecycle Task Packages, Build-MemSecBench-Case Skill, Lifecycle Evaluation Workflow, Evidence-Based LLM Judging, and a 24-configuration matrix to trace malicious semantics through memory systems. MemSecBench reports 50.3% End-to-End Attack Success Rate and 56.1% Selective Repair Success Rate across 310 cases and 24 configurations, contrasting Native with Mem0, Mem0-Graph, and A-MEM.

Long-Term Memory

MemTrace: Probing What Final Accuracy Misses in Long-Term Memory

Xianxuan Long, Zhikai Chen et al.

arXiv 2026 · 2026

MemTrace builds typed Knowledge points, Probe construction, Metrics, and Diagnostic views to trace each user fact across sessions and query styles. MemTrace’s main finding is that across 13 configurations on 835 knowledge points and 15,422 question rows, failures arise about 10× more often from unused reachable evidence than from missing evidence.

Benchmark

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Mengru Wang, Haozhe Luo et al.

arXiv 2026 · 2026

MemTrapBench combines a Taxonomy, Instance Construction, and Two-Gate Quality Flow to generate 1,050 adversarial multi-turn dialogues that trigger Reasoning Fixation and Belief Distortion in LLM memory use. On Gemini-3-Flash-Preview with LightMem, AdaptiveMem recovers 14.9 percentage points on MemTrapBench while maintaining performance on LongMemEval.

Memory Architecture

Mental Model Management: An Operator-Based Framework for LLM Memory

Oliver Kramer

arXiv 2026 · 2026

Mental Model Management (3M) represents knowledge as evolving Mental Models, each composed of compact Chunks, and transforms them using operators like Extract, Add, Update, Merge, and Abstract. In a cold-start Evolution Strategies ingestion, 3M reduced 1,577 input words to 1,060 words across three linked models while creating 24 relations and four knowledge gaps.

Long-Term Memory

MobileMem: Learning from a Year of Mobile Experiences

Xinle Deng, Yida Xue et al.

arXiv 2026 · 2026

MobileMem builds long‑horizon mobile interaction trajectories using components like User Prior Knowledge Construction, KEME, User Trajectory Synthesis, QA Pair Synthesis, and Quality Control to stress on‑device memory layers. On the MobileMem benchmark, A‑MEM and HippoRAG2 reach overall LLM‑Judge scores up to 80.06% with GPT‑5.4‑mini, compared to 45.19% for Long Context, while revealing token costs as high as 11,170.44k tokens per trajectory.

Agent Memory

Multi-Agent Memory from a Computer Architecture Perspective: Visions and Challenges Ahead

Zhongming Yu, Naicheng Yu et al.

arXiv 2026 · 2026

Multi-Agent Memory Architecture organizes Agent IO Layer, Agent Cache Layer, Agent Memory Layer, Agent Cache Sharing, and Agent Memory Access Protocol into a computer-architecture-style design for LLM agents. Multi-Agent Memory Architecture’s main result is a conceptual unification of shared and distributed memory plus a research agenda for multi-agent memory consistency instead of benchmark gains.

Memory Architecture

MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

Huawei Lin, Peng Li et al.

arXiv 2026 · 2026

MUSE-Autoskill combines a Master Agent, Skill Creator, Skill Bank, Evaluator, and Memory (short-term, long-term, skill-level) into a unified skill lifecycle that creates, executes, and refines skills in-context. On the 75-task SkillsBench common set, MUSE-Autoskill achieves 59.67% accuracy with human skills, +12.72 percentage points over its no-skill setting and ahead of Hermes, Codex, and Claude Code.