Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 1 of 19

Agent Memory

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Xiaoyang Li, Yiqi Wang et al.

arXiv 2026 · 2026

Correlated Promotion Benchmark (CPB) combines CPB-Static, CPB-Live, a gold admission rule, lineage collapse, and a governance rule to stress-test epistemic admission in shared agent memory. On CPB-Live, the governance rule keeps damage shares between 0.112 and 0.152 and false adoption between 0.06 and 0.09, while majority vote and LLM judges often match share-all’s false adoption.

Benchmark

According to Me: Long-Term Personalized Referential Memory QA

Jingbiao Mei, Jinghong Chen et al.

arXiv 2026 · 2026

ATM-Bench structurally evaluates long-term multimodal personal memory using Memory Ingestion, Retrieval, and Answer Generation with Schema-Guided Memory and Descriptive Memory variants. On ATM-Bench-Hard, Oracle with SGM reaches 47.3% QS while the best full system stays under 20% accuracy, revealing a large gap.

Memory Architecture

A Control Architecture for Training-Free Memory Use

Yanzhen Lu, Muchen Jiang et al.

· 2026

TAG routes low-confidence steps to uncertainty-based routing, filters them with guarded acceptance with rollback, chooses between bank selection across rule and exemplar memory, and prunes via evidence-based retirement inside a unified control loop. On SVAMP and ASDiv, TAG reaches 81.0% and 85.2% accuracy, improving over the 74.0% and 77.5% no-memory baselines while a compute-matched Retry baseline stays flat.

BenchmarkAgent Memory

Active Context Compression: Autonomous Memory Management in LLM Agents

Nikhil Verma

· 2026

Focus Agent adds start_focus, complete_focus, a persistent Knowledge block, and an optimized Persistent Bash plus String-Replace Editor scaffold to actively compress context during long software-engineering tasks. On five hard SWE-bench Lite instances against a Baseline ReAct agent, Focus Agent achieves 22.7% token reduction (14.9M → 11.5M) while matching 3/5 = 60% task success.

Memory Architecture

ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents

Song-Li Wu, Jingyi Wang et al.

arXiv 2026 · 2026

ActiveMem organizes experiences with a Hierarchical Latent Memory Tree, Dual-Head Memory Controller, Latent Injection Head, and Tree Action Head to build dependency-aware memory paths. On ALFWorld with Qwen3-8B, ActiveMemGRPO reaches 95.57% vs MemGenGRPO’s 90.60%, while also boosting TriviaQA from 80.65% to 87.46%.

Agent Memory

ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents

Xiaohui Zhang, Zequn Sun et al.

· 2026

ActMem transforms dialogue history into atomic facts via Memory Fact Extraction, groups them with Fact Clustering, links them through a Memory KG Construction module, and uses Counterfactual-based Retrieval and Reasoning for action-aware answers. On ActMemEval, ActMem reaches 76.52% QA accuracy with DeepSeek-V3, beating LightMem’s 63.97% by 12.55 points and NaiveRAG’s 61.54%.

Agent Memory

AdaMEM: Test-Time Adaptive Memory for Language Agents

Yunxiang Zhang, Yiheng Li et al.

arXiv 2026 · 2026

AdaMEM combines Long-Term Trajectory Memory, Short-Term Strategy Memory, ADAMEM-HIGH, ADAMEM-LOW, and STEP-MFT to adapt agent behavior at each decision step without parameter updates. On ALFWorld unseen, AdaMEM-LOW achieves 58.2% Success Rate versus 52.2% for Synapse, and on WebShop reaches 74.2 Task Score versus 71.4 for the No Memory baseline.

Agent MemoryLong-Term Memory

Adaptive Memory Admission Control for LLM Agents

Guilin Zhang, Wei Jiang et al.

· 2026

A-MAC scores candidate memories using Utility, Confidence, Novelty, Recency, and Type Prior combined by a learned linear admission policy with Algorithm 1 A-MAC Memory Admission. On the LoCoMo benchmark, A-MAC achieves F1 0.583 and 2644 ms latency, improving F1 by 0.042 and reducing latency by 1187 ms compared to A-mem.

Benchmark

AdMem: Advanced Memory for Task-solving Agents

Runzhe Wang, Huilin Lu et al.

arXiv 2026 · 2026

AdMem unifies Actor agent, Long-term memory agent, Critic agent, Short-term memory, and Reward-based long-term memory management into a bi-level semantic episodic procedural memory system for task-solving agents. On AgentBoard’s AlfWorld domain, AdMem reaches 63.4% task completeness versus 49.3% for ReAct, a +14.1 percentage point gain.

Long-Term Memory

Advancing Open-source World Models

Robbyant Team, Zelin Gao et al.

arXiv 2026 · 2026

LingBot-World combines a Data Engine, Fundamental World Model, Action-Conditioned World Model, and Post-Training causal adaptation to turn a 28B-parameter video generator into a real-time interactive world simulator. On the VBench benchmark, LingBot-World achieves a dynamic degree of 0.8857 versus 0.7612 for Yume-1.5, while also improving imaging quality to 0.6683.

RAG

A Dynamic Retrieval-Augmented Generation System with Selective Memory and Remembrance

Okan Bursa

· 2026

Adaptive RAG Memory (ARM) augments a standard retriever–generator stack with a Dynamic Embedding Layer and Remembrance Engine that track usage statistics and apply selective remembrance and decay to embeddings. On a lightweight retrieval benchmark, ARM achieves NDCG@5 ≈ 0.9401 and Recall@5 = 1.000 with 22M parameters, matching larger baselines like gte-small while providing the best efficiency among ultra-efficient models.

Cognitive ArchitectureAgent Memory

Aeon: High-Performance Neuro-Symbolic Memory Management for Long-Horizon LLM Agents

Mustafa Arslan

· 2026

Aeon restructures LLM memory using the Atlas, Trace, Semantic Lookaside Buffer, Write Ahead Log, and Sidecar Blob Arena inside a zero copy Core Shell kernel. Aeon achieves 4.70 ns INT8 dot products, 3.09 µs Atlas traversal at 100K nodes, 3.1× compression, and P99 read latency of 750 ns under 16 thread contention compared to FP32 and flat scan baselines.

Benchmark

AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

Yiheng Shu, Bernal Jiménez Gutiérrez et al.

arXiv 2026 · 2026

AGENTCL evaluates continual learning in language agents by contrasting naive and compositional task streams with a two-pass protocol and MEMPROBE’s interaction memory, insight memory, and skill memory. On the CodeEval-Pro compositional stream, AGENTCL shows MEMPROBE achieving 66.7% first-pass accuracy versus the memoryless ReAct’s 48.8% (+17.9 pp), while naive streams show much smaller separations.

BenchmarkBenchmarkLong-Term Memory

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

Manoj Madushanka Perera, Adnan Mahmood et al.

· 2026

AgenticAI-DialogGen chains ChatPreprocessor, KnowledgeExtractor, TopicAnalyzer, KnowledgeGraphBuilder, PersonaGenerator, DuelingChat Agent, ConversationValidator, ConversationRefiner, QAGeneration, and PostProcessing to turn raw multi-session chats into topic-guided, persona-grounded conversations with explicit short- and long-term memories. On the TGC / KG memory QA benchmark, Mistral-7B fine-tuned within AgenticAI-DialogGen achieves 87.36 F1, compared to GPT-4’s 83.77 F1 in a zero-shot setting on the same task.

BenchmarkAgent Memory

Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

Yi Yu, Liuyi Yao et al.

arXiv 2026 · 2026

Agentic Memory (AgeMem) exposes memory management tools, a three-stage progressive RL strategy, and step-wise GRPO directly inside the agent policy to jointly control long-term and short-term memory. On Qwen3-4B-Instruct, AgeMem attains 54.31% average performance across ALFWorld, SciWorld, PDDL, BabyAI, and HotpotQA, exceeding the best baseline A-Mem at 45.74%.

Benchmark

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Xiangchen Cheng, Yunwei Jiang et al.

arXiv 2026 · 2026

AgenticSTS composes each decision from L1 protocol instructions, L2 state-typed prompts, L3 game knowledge, L4 episodic memory, and L5 skill library instead of an accumulating transcript. On fixed A0 Slay the Spire 2 runs, AgenticSTS with L5 skills wins 6/10 games versus 3/10 for the no-scaffold baseline, and climbs to A6–A8 in auto-mode streams.

Memory Architecture

AgentIR: A Workload-Adaptive Cascade Retrieval Substrate for Long-Term Conversational Memory

Aojie Yuan, Haiyue Zhang, Shahin Nazarian

arXiv 2026 · 2026

AgentIR combines a SIMD-Accelerated BM25 engine, a Temporal Partitioned Index, an Agent-Aware Fusion layer, and a Cascade Router over a shared CSR substrate to adapt retrieval per query. On LongMemEval, AgentIR’s confidence-triggered cascade skips dense on 63% of queries for a 2.67× latency reduction at parity LLM-judged accuracy, and on BEIR AgentIR’s BM25 core is up to 29× faster than Pyserini 8T at matching nDCG@10.

Benchmark

AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents

Ahmed Cherif

arXiv 2026 · 2026

AgentMemBench wires five strategies—In-Context Windowing (ICW), External Key-Value Store (EKV), Graph-Based Episodic Memory (GEM), Compression-Based Summarisation (CBS), and Web-Augmented Memory (WAM)—into a shared Store Read Generate interface over multi-session dialogues. AgentMemBench’s main result is that EKV achieves macro Recall@5 0.792 on LoCoMo, MultiDoc2Dial, and MSC, a +0.324 gain over ICW’s 0.468, while also leading MRR, Answer F1, and Faithfulness.

BenchmarkAgent Memory

Agent Memory Below the Prompt: Persistent Q4 KV Cache for Multi-Agent LLM Inference on Edge Devices

Yakov Pyotr Shkolnikov

· 2026

Agent Memory Below the Prompt stores each agent’s KV state in a block pool, quantizes it via a Q4 pipeline, reloads it with BatchQuantizedKVCache, and reuses it across phases using cross-phase context injection. On Gemma 3 12B, Agent Memory Below the Prompt reduces cold TTFT from 172,096 ms to 1,264 ms at 32K context (136×) compared to FP16 prefix caching baselines like vllm-mlx.