Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 2 of 19

Agent Memory

Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads

Yasmine Omri, Ziyu Gan et al.

arXiv 2026 · 2026

Agent Memory decomposes agent workloads into ingestion, memory construction, storage, retrieval, prompt assembly, generation, and maintenance, and classifies ten systems across four paradigms. Agent Memory’s profiling on MemoryAgentBench and MemoryArena reveals over 47× spread in lifecycle energy per correct answer and two orders of magnitude differences in serving latency across paradigms.

Agent Memory

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

Taeil Kim, Kangsan Kim, Sung Ju Hwang

arXiv 2026 · 2026

Agent Memory Distillation builds Workflow memory, Subtask memory, and Function memory from successful teacher trajectories and injects them via proactive and reactive retrieval to guide small agents. On AppWorld, Agent Memory Distillation lifts Qwen3-4B from 14.88% to 49.40% accuracy (+34.52%p) compared to the zero-shot baseline.

Agent Memory

Agent Memory Is a Surface for Endogenous Authorization Laundering

Tommaso Cerruti, Mika Okamoto, Ansel Kaplan Erol

arXiv 2026 · 2026

EAL-BENCH wires a memory writer, executor, canonical ledger, and authorization predicate together to isolate how agent memory encodes and propagates permissions. Across procurement, cybersecurity, and finance, typed incremental memory forms false authority for up to 50.2% of unauthorized requests, and executors act on it in 98.6% of matched trials.

Benchmark

AgentSM: Semantic Memory for Agentic Text-to-SQL

Asim Biswal, Chuan Lei et al.

· 2026

AgentSM builds a planner agent, schema linking agent, trajectory store, and composite tools that turn past execution traces into reusable, structured semantic memory. On Spider 2.0 Lite, AgentSM with Claude 4 Sonnet achieves 44.8% execution accuracy versus 28.7% for SpiderAgent, while shortening trajectories and latency.

Benchmark

AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems

Zachary Johnson, Nigel Boachie Kumankumah et al.

arXiv 2026 · 2026

AIM orchestrates an Operation Detection Agent, Privacy-Aware Classification Engine, Structured Memory Extraction, In-Context Semantic Deduplication, and Cascaded Memory Retrieval over a shared vector-indexed store. On MUMBench, AIM with GPT-5.4-mini achieves 58.8% strict and 70.5% state-aware operation accuracy while reaching 96.0% visibility classification accuracy with GPT-5.4.

Cognitive ArchitectureAgent Memory

Aligning Progress and Feasibility: A Neuro-Symbolic Dual Memory Framework for Long-Horizon LLM Agents

Bin Wen, Ruoxuan Zhang et al.

· 2026

Neuro-Symbolic Dual Memory Framework uses Progress Memory, Feasibility Memory, a Blueprint Planner Agent, a Progress Monitor Agent, and an Actor Agent to decouple semantic progress guidance from executable feasibility checks. On ALFWorld, Neuro-Symbolic Dual Memory Framework achieves 94.78% success rate versus 88.81% for AWM, and on WebShop reaches 0.7132 score versus 0.5998 for WALL-E 2.0.

Benchmark

AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment

Jianfei Xiao, Xiang Yu et al.

· 2026

AlpsBench combines Personalized Information Extraction, Personalized Information Update, Personalized Information Retrieval, and Personalized Information Utilization over 2,500 WildChat dialogues with human-verified structured memories. AlpsBench shows, for example, that Gemini-3 Flash scores 51.67 on Task 1 Extraction while DeepSeek Reasoner reaches 0.9569 retrieval recall with 100 distractors on AlpsBench.

Agent MemoryLong-Term Memory

AMA: Adaptive Memory via Multi-Agent Collaboration

Weiquan Huang, Zixuan Wang et al.

· 2026

AMA orchestrates four agents — the Constructor, Retriever, Judge, and Refresher — to build Raw Text, Fact Knowledge, and Episode Memory and route queries adaptively across these granularities. On the LoCoMo benchmark with GPT-4.1-mini, AMA achieves an overall LLM Score of 0.805 compared to Nemori’s 0.774, while reducing token consumption by approximately 80% relative to FullContext.

BenchmarkBenchmarkLong-Term Memory

A-MBER: Affective Memory Benchmark for Emotion Recognition

Deliang Wen, Ke Sun, Yu Wang

· 2026

A-MBER builds multi-session conversational scenarios via a staged pipeline of persona specification, long-horizon planning, conversation generation, annotation, question construction, and benchmark-unit packaging. On A-MBER, a structured memory system reaches 0.69 judgment accuracy, 0.66 retrieval, and 0.65 explanation versus 0.34, 0.29, and 0.31 for a no-memory baseline.

BenchmarkAgent Memory

AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations

Cheng Jiayang, Dongyu Ru et al.

· 2026

AMemGym combines Structured Data Generation, On-Policy Interaction, Evaluation Metrics, and Meta-Evaluation to script user state trajectories, drive LLM-simulated role-play, and score write–read–utilization behavior. On AMemGym’s base configuration, AWE-(2,4,30) reaches a 0.291 normalized memory score on interactive evaluation, while native gpt-4.1-mini only achieves 0.203, exposing substantial gaps between memory agents and plain long-context LLMs.

Agent MemoryLong-Term Memory

AMV-L: Lifecycle-Managed Agent Memory for Tail-Latency Control in Long-Running LLM Systems

Emmanuel Bamidele

· 2026

AMV-L manages agent memory using a Memory Value Model, Tiered Lifecycle, Bounded Retrieval Path, and Lifecycle Manager to decouple retention from retrieval eligibility. Under a 70k-request long-running workload, AMV-L improves throughput from 9.027 to 36.977 req/s over TTL and reduces p99 latency from 5398.167 ms to 1233.430 ms while matching LRU’s retrieval quality.

SurveyAgent Memory

Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations

Dongming Jiang, Yi Li et al.

arXiv 2026 · 2026

Anatomy of Agentic Memory organizes agentic memory into four structures using components like Lightweight Semantic Memory, Entity-Centric and Personalized Memory, Episodic and Reflective Memory, and Structured and Hierarchical Memory. Anatomy of Agentic Memory then reports comparative results such as Nemori’s 0.781 semantic judge score on LoCoMo versus SimpleMem’s 0.298, and latency differences like 1.129s for Nemori versus 32.372s for MemoryOS.

BenchmarkBenchmark

APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay

Pratyay Banerjee, Masud Moshtaghi, Ankit Chadha

· 2026

APEX-EM combines a Procedural Knowledge Graph, Experience Memory store, PRGII workflow, Task Verifiers, and StructuralSignatureExtractor to store and reuse full procedural-episodic traces without changing model weights. On KGQAGen-10k, APEX-EM reaches 89.6% accuracy (95.3% CSR) versus 41.3% without memory and surpasses the GPT-4o w/ SP oracle at 84.9%.

Long-Term Memory

ArborMem: Navigating Interaction States with Memory Forests

Zongwei Lv, Yuemeng Xu et al.

arXiv 2026 · 2026

ArborMem organizes conversations into a Conversation forest, Reusable evidence store, State localization, Context assembly and generation, and Online memory commit to track resumable trajectories. On BranchMemEval, ArborMem achieves 81.00% accuracy, beating Mem0 at 76.00% by 5.00 percentage points.

RAG

Are We Ready For An Agent-Native Memory System?

Wei Zhou, Xuanhe Zhou et al.

arXiv 2026 · 2026

Are We Ready For An Agent-Native Memory System? analyzes Memory Representation and Storage, Memory Extraction, Memory Retrieval and Routing, and Memory Maintenance across 12 real systems like Mem0, Zep, MemTree, LightMem, MemOS, MemoryOS, and A-MEM. The study’s main result is that structured systems such as Zep reach 48.0 LLM Judge Accuracy on LongMemEval while Long Context reaches 19.0, and that localized maintenance strategies like LightMem achieve 48.3% normalized utility at only 3.67 s per query.

Benchmark

ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs

Jianlong Lei, Shashikant Ilager

· 2026

ARKV dynamically combines Per-layer OQ ratio estimation, Token importance scoring, and Tri-state cache assignment to manage KV cache precision under a global memory budget. On LongBench, ARKV reaches 0.972 relative performance versus 0.979 for Origin while achieving 4× KV memory reduction and maintaining ~86% Tokens Per Second.

BenchmarkBenchmark

Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents

Yuxuan Cai, Jie Zhou et al.

· 2026

PROACTAGENT combines Experience-Enhanced Online Evolution (EXPONEVO), a structured EXPERIENCE BASE, and Proactive Reinforcement Learning-based Retrieval (PROACTRL) to jointly evolve memory and policy with retrieval as an explicit action. On SciWorld, PROACTAGENT reaches 73.50% SR versus 55.50% for GRPO+Reflexion, while cutting interaction rounds from 27.52 to 18.38.

SurveyBenchmarkAgent MemoryLong-Term MemoryMemory Architecture

A Survey on the Security of Long-Term Memory in LLM Agents: Toward Mnemonic Sovereignty

Zehao Lin, Chunyu Li, Kai Chen

· 2026

Mnemonic Sovereignty analyzes long term Write, Store, Retrieve, Execute, Share, and Forget Rollback phases against integrity, confidentiality, availability, and governance objectives for agent memory. Mnemonic Sovereignty’s lifecycle matrix shows most of the ~70 works cluster on write and retrieve integrity, leaving store, availability, and governance primitives like write gate validation and post deletion verification almost entirely unexplored.

BenchmarkAgent Memory

ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks

Samuel Sameer Tanguturi

· 2026

ATANT v1.1 structurally analyzes seven benchmarks using the 7 v1.0 continuity properties, the 10 checkpoints, a property-coverage matrix, and the Kenotic v1.0 reference implementation. ATANT v1.1 reports 96% ATANT cumulative-scale versus 8.8% LOCOMO substring accuracy, showing that LOCOMO, LongMemEval, BEAM, MemoryBench, Zep eval, MemGPT/Letta, and RULER measure different properties from continuity.

Agent Memory

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory

Zitong Shi, Yixuan Tang, Anthony Kum Hoe Tung

arXiv 2026 · 2026

A-TMA wraps existing memory systems with Bank Level State Maintenance, Retrieve Level Evidence Construction, and QA Level Evidence State Conditioning to track current, historical, and transition facts explicitly. On the LTP benchmark, A-TMA lifts Graphiti/Zep conflict accuracy from 0.480 to 0.720 and temporal F1 on LoCoMo from 0.0295 to 0.1705 over Graphiti/Zep.