Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 10 of 19

Long-Term Memory

MemTrace: Probing What Final Accuracy Misses in Long-Term Memory

Xianxuan Long, Zhikai Chen et al.

arXiv 2026 · 2026

MemTrace builds typed Knowledge points, Probe construction, Metrics, and Diagnostic views to trace each user fact across sessions and query styles. MemTrace’s main finding is that across 13 configurations on 835 knowledge points and 15,422 question rows, failures arise about 10× more often from unused reachable evidence than from missing evidence.

Benchmark

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Mengru Wang, Haozhe Luo et al.

arXiv 2026 · 2026

MemTrapBench combines a Taxonomy, Instance Construction, and Two-Gate Quality Flow to generate 1,050 adversarial multi-turn dialogues that trigger Reasoning Fixation and Belief Distortion in LLM memory use. On Gemini-3-Flash-Preview with LightMem, AdaptiveMem recovers 14.9 percentage points on MemTrapBench while maintaining performance on LongMemEval.

Benchmark

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

Ryuichi Sumida, Koji Inoue, Tatsuya Kawahara

arXiv 2026 · 2026

MemUse combines a 4-month GPT-4.1-mini deployment with components like Summary-only, LC-k%, RAG-k%, and the MEMUSE benchmark to study memory use in real conversations. MemUse’s main result shows Direct QA rising from 19.7% to 70.1% across conditions while satisfaction stays flat, and within MEMUSE moments Natural Integration, not Direct QA, is associated with higher user satisfaction.

Memory Architecture

Mental Model Management: An Operator-Based Framework for LLM Memory

Oliver Kramer

arXiv 2026 · 2026

Mental Model Management (3M) represents knowledge as evolving Mental Models, each composed of compact Chunks, and transforms them using operators like Extract, Add, Update, Merge, and Abstract. In a cold-start Evolution Strategies ingestion, 3M reduced 1,577 input words to 1,060 words across three linked models while creating 24 relations and four knowledge gaps.

Benchmark

MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents

Kaichao Liang, Yuqi Cui et al.

arXiv 2026 · 2026

MindMemOS organizes memories in a unified entity–property–time graph, with MindVanilla, MindSchema, MindMemEvolve, dreaming, feedback, and MindSkillEvolve coordinating memory and skill evolution. MindMemOS reaches 94.03% overall accuracy on LOCOMO and 70.63% on PersonaMem, surpassing EverOS by 0.98 and 3.06 percentage points respectively.

Long-Term Memory

MobileMem: Learning from a Year of Mobile Experiences

Xinle Deng, Yida Xue et al.

arXiv 2026 · 2026

MobileMem builds long‑horizon mobile interaction trajectories using components like User Prior Knowledge Construction, KEME, User Trajectory Synthesis, QA Pair Synthesis, and Quality Control to stress on‑device memory layers. On the MobileMem benchmark, A‑MEM and HippoRAG2 reach overall LLM‑Judge scores up to 80.06% with GPT‑5.4‑mini, compared to 45.19% for Long Context, while revealing token costs as high as 11,170.44k tokens per trajectory.

Agent Memory

Multi-Agent Memory from a Computer Architecture Perspective: Visions and Challenges Ahead

Zhongming Yu, Naicheng Yu et al.

arXiv 2026 · 2026

Multi-Agent Memory Architecture organizes Agent IO Layer, Agent Cache Layer, Agent Memory Layer, Agent Cache Sharing, and Agent Memory Access Protocol into a computer-architecture-style design for LLM agents. Multi-Agent Memory Architecture’s main result is a conceptual unification of shared and distributed memory plus a research agenda for multi-agent memory consistency instead of benchmark gains.

Memory Architecture

MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

Huawei Lin, Peng Li et al.

arXiv 2026 · 2026

MUSE-Autoskill combines a Master Agent, Skill Creator, Skill Bank, Evaluator, and Memory (short-term, long-term, skill-level) into a unified skill lifecycle that creates, executes, and refines skills in-context. On the 75-task SkillsBench common set, MUSE-Autoskill achieves 59.67% accuracy with human skills, +12.72 percentage points over its no-skill setting and ahead of Hermes, Codex, and Claude Code.

Agent Memory

MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence

Walid Saidi

arXiv 2026 · 2026

MutMem V2 wires a portable recall-disclosure envelope, cross-object predicates, portable mutation profile, and independent Node/Python verifiers into a single evidence protocol for persistent agent memory. MutMem V2’s main result is exact 72/72 parity on terminal verdict and primary reason between Node and Python verifiers, plus 42/42 agreement on a production-derived conformance corpus.

Memory Architecture

OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents

Yulin Hu, Zimo Long et al.

· 2026

OP-Bench benchmarks over-personalization using Irrelevance, Repetition, and Sycophancy across 1,700 long-horizon queries built from LoCoMo, then analyzes memory systems like RAG, LDAgent, Mem0, MemU, and MEMOS. Self-ReCheck, a lightweight memory filter, improves average OP-Bench scores by 29% for Qwen3-8B while preserving personalization on LoCoMo compared to BASE.

Long-Term Memory

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

Bowen Yang, Kaiming Jin et al.

arXiv 2026 · 2026

OS-SYMPHONY coordinates an Orchestrator, Reflection-Memory Agent, and Versatile Tool Agents (Multimodal Searcher, Grounders, Coder) to stabilize long-horizon GUI workflows and fetch visual tutorials on demand. On OSWorld-Verified, OS-SYMPHONY with GPT-5 scores 65.84% at 100 steps, beating Agent S3 w/ GPT-5 (62.63%) by 3.21 percentage points.

Benchmark

PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use

Mingfei Lu, Mengjia Wu, Yi Zhang

arXiv 2026 · 2026

PairPref constructs controlled pairs using Stored preferences, Companion memory pool, Multi-stage validation, and a dual Selection and Free generation protocol to test when preferences should guide answers. On the 1,227-pair benchmark, PairPref reports selection selectivity up to 64.9 points for Claude Opus 5 but free-generation Pair success of only 3.6–18.3% across eight models.

Long-Term Memory

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory

Zhifei Xie, Zongzheng Hu et al.

· 2026

Pask combines Demand Detection (IntentFlow), Pask-MM hierarchical memory, and the Pask-PAS proactive agent system to continuously infer latent needs and act through tools and frontier models. On LatentNeeds-Bench, Pask’s IntentFlow achieves 84.2% balanced accuracy versus 80.8% for Gemini-3-Flash, under ~1.3–1.5 s per-turn latency.

Benchmark

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

Shuhan Xue, Zixin Ding et al.

arXiv 2026 · 2026

PAST-Bench evaluates recursive self-improvement in personal agents by running ordered task-family sequences with Cold, Learn, Evaluation, and Control episodes across four capabilities, and by instrumenting Hermes+ with Plan, Render, Route, Gate, and Close mechanisms. Hermes+ on MiniMax-M2.7 lifts the family-balanced Overall Δ from +0.13 to +0.15 and the mechanism-evidence score from 0.64 to 0.73 on PAST-Bench, with a +0.12 gain on Update (from +0.12 to +0.24) over the Hermes baseline.

Long-Term Memory

PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments

Shuochen Liu, Junyi Zhu et al.

· 2026

PERMA constructs temporally ordered Interaction Events, a dynamic Persona State, and a dual-process Memory System over Clean, In-session Noise, and Style-aligned Long-context datasets. PERMA shows that structured memory agents like MemOS can compress raw 34k-token histories down to ~700 tokens while preserving higher MCQ accuracy than vanilla RAG in realistic, noisy multi-session personalization tasks.

Long-Term Memory

PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?

Sidharth Pulipaka, Oliver Chen et al.

· 2026

PersistBench combines Monte Carlo Tree Search, Seed Initialization and Candidate Generation, Search and Scoring, and Human Verification to build realistic long-term memory test cases across cross-domain leakage, sycophancy, and beneficial memory use. PersistBench then reports a median 53% cross-domain leakage failure rate and 97.8% sycophancy failure rate on 18 LLMs using its 500-sample benchmark.

Memory Architecture

PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents

Hanzhong Zhang, Ziwei Xiang et al.

arXiv 2026 · 2026

PersMem maps a fixed personality vector into Affective Appraisal, Memory Retention, Passive Affect-Driven Memory Retrieval, and Active Goal-Driven Memory Retrieval to control how autobiographical memories are stored and recalled. On attachment and Big Five evaluations, PersMem reaches 48.1% four-way attachment classification and 67.5% Big Five dialogue discrimination, exceeding chance and no-trait controls.

Benchmark

Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents

Yeonjun In, Wonjoong Kim et al.

arXiv 2026 · 2026

Personalize-then-Store combines PerMem-Bench, PerMem-Benchs, PerMem-Benchd, session-level storage gating, and structural note modeling to study personalized memory for long-horizon agents. On PerMem-Bench, Oracle gating substantially boosts Memory Retention Rate over Universal policies, especially at 100–200 entry budgets, while existing gating baselines like Greedy and Structure-aware recover only incremental gains.

Benchmark

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

Seungbin Yang, Chaewoon Ki et al.

arXiv 2026 · 2026

PACMEM in PERSONATRAIL builds Factual Memory, Preference Memory, Trajectory Segmentation, and Memory Retrieval to structure browser-level histories for personalized web navigation. On PERSONATRAIL, PACMEM achieves 53.61% Task Success Rate on single-hop preference inference with Qwen3.6-27B, compared to 46.45% for ReasoningBank, and 39.03% vs 17.84% on multi-hop episodic grounding.

Long-Term Memory

PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents

Ke Yang, Zixi Chen et al.

· 2026

PLUGMEM standardizes episodic traces with the Structuring Module, retrieves via an abstraction-aware Retrieval Module, and compresses outputs through a Reasoning Module into a unified memory graph. On HotpotQA, PLUGMEM reaches 61.4 EM and 74.1 F1 with only 81.6 memory tokens, compared to 51.7 EM and 62.7 F1 with 659.2 tokens for Vanilla Retrieval.

Agent Memory

Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents

Wei Zou, Mingwen Dong et al.

· 2026

eTAMP stores attacker-crafted payloads inside Trajectory Memory, then reactivates them via semantic Cross-Site Task Pairing and Chaos Monkey induced frustration during later tasks. On (Visual)WebArena, eTAMP achieves up to 32.5% ASRB on GPT-5-mini and 23.4% on GPT-5.2 under Frustration Exploitation, showing persistent cross-session compromise without direct memory access.

Benchmark

QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents

Heng Wang, Yifei Li et al.

arXiv 2026 · 2026

QUMem combines Dynamic Episode Construction, Typed Memory Decomposition, and Query-Conditioned User-State Inference to segment histories into semantically coherent episodes and split them into factual, preference, and transferable insight memories. On PersonaMem, QUMem reaches 70.58% overall accuracy with Gemini-3.5-flash, improving over Mem0’s 63.29% and Zep’s 54.46%.