Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 7 of 19

Memory Architecture

InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain

Tiancheng Han, Yong Li et al.

arXiv 2026 · 2026

InfoMem trains chunk-wise memory agents by comparing per-token log-likelihood of the ground-truth answer with and without the final memory under Group Relative Policy Optimization and a final-memory reward definition. InfoMem reaches 19.453% on CorpusQA versus 16.413% for Outcome-only GRPO and 1.520% for ReMemR1, while also improving LongMemEval, MRCR-8needle, and RULER synthetic QA.

Agent Memory

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

Hanling Tian, Gengyu Zhang et al.

arXiv 2026 · 2026

InjecMEM combines a retriever agnostic anchor, adversarial command, Multi GCG, and the MemoryOS hierarchy of STM, MTM, and LPM to inject a single poisoned interaction that later steers topic specific queries. On MemoryOS health finance and agriculture domains, InjecMEM achieves up to 35.4% RSR@50 and 76.6% ASR c, surpassing centroid plus GCG and other baselines.

Agent Memory

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Yefan Zhou, Yang Li et al.

arXiv 2026 · 2026

Just-in-Time Memory (JITMEM) combines a Memory Bank, Retriever, Memory Curator, and frozen Agent Executor to store raw trajectories and distill task-conditioned payloads only when needed. On WebShop, JITMEM with Qwen3-8B curator and Gemini-2.5-Pro executor reaches 50.5% success rate, improving over RL-trained SkillOS at 41.3% by +9.2 points.

Agent Memory

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Ruizhe Li, Mingxuan Du et al.

arXiv 2026 · 2026

Keep It InMind introduces the InMind benchmark plus paired controls (Naive query, Indirect query, Target recall, Backbone control) to isolate failures in agent memory use. On InMind, Keep It InMind finds that the backbone reaches 84.0% indirect accuracy with the memory in context, while six retrieval-based systems reach at most 14.4% application accuracy.

Agent Memory

LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory

Jing Yu, Yibo Zhao et al.

arXiv 2026 · 2026

LazyMem stores raw histories and at query time runs Hybrid Retrieval, History Windowing, and a trained Memory-Processing Model to build compact, query-conditioned memory for an answer LLM. On LongMemEval, LazyMem-4B achieves 0.85 LLM-judge accuracy with 213 answer-context memory tokens, a 0.03 gain over StructMem while using 21.0× fewer tokens.

Long-Term Memory

LeanMem: Simple and Efficient Long-Term Memory for LLM Agents

Yuxin Liao, Le Wu et al.

arXiv 2026 · 2026

LeanMem combines Controlled Memory Writing, Selective Memory Evolution, and Adaptive Evidence Composition to route dialogue into profile, event, and record memories with minimal tokens. On LongMemEval-S with GPT-4.1-mini, LeanMem achieves 91.80% Accuracy and 97.67% Recall, improving A-Mem by 15.07 Accuracy points while cutting construction tokens from 1330.77K to 117.61K.

BenchmarkBenchmarkCognitive Architecture

Learning to Forget: Sleep-Inspired Memory Consolidation for Resolving Proactive Interference in Large Language Models

Ying Xie

· 2026

SleepGate augments transformers with a Conflict-Aware Temporal Tagger, Forgetting Gate, Consolidation Module, and Sleep Trigger that periodically rewrite the KV cache during sleep micro-cycles. On the PI-LLM benchmark, SleepGate achieves 99.5% retrieval accuracy at PI depth 5 and 97.0% at depth 10, while full KV cache, sliding window, H2O, StreamingLLM, and a decay-only ablation all stay below 18% across all depths.

Agent Memory

Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems

Rakibul Hasan Rajib, Mengxing Zheng, Qian Lou

arXiv 2026 · 2026

Gated-Memory Routing coordinates a multi-agent LLM system through a History-Aware Role Allocator, LLM Router, Retrieval Gate, Memory Write Gate, and Adaptive Halting Controller over a gated execution memory. On five benchmarks, Gated-Memory Routing achieves 77.73% average accuracy on MATH, GSM-Hard, MBPP, HumanEval, and MMLU-Pro, improving by 2.44 points over Puppeteer (qwen-2.5-32B) while reducing HumanEval inference cost by 31.9%.

Long-Term Memory

LifeBench: A Benchmark for Long-Horizon Multi-Source Memory

Zihao Cheng, Weixin Wang et al.

arXiv 2026 · 2026

LifeBench simulates long-horizon personal trajectories via Persona Synthesis, Hierarchical Outline Planning, Dual-Agent Daily Activity Simulation, Phone Data Generation, and Question Answering Generation. On LifeBench’s 2,003-question benchmark, MemOS achieves 55.22% overall accuracy, far below its ~90% on LoCoMo and LongMemEval, revealing substantial room for progress.

Long-Term Memory

LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

Hanyu Zhao, Yuqian Feng et al.

arXiv 2026 · 2026

LifeFuse-Mem separates memory into Lifecycle-Aware Subspace Partition, Permanent Routing and Subspace Writing, and Protected State Fusion and Lifecycle Readout to keep permanent and temporary information from colliding. On the Hard Attribution Anti-Overwrite benchmark, LifeFuse-Mem raises acquisition-controlled retention to 63.71% and 69.39% over δ-Mem while reducing overwrite to 36.29% and 30.61%.

Agent Memory

LifeSide: Benchmarking Agents as Lifelong Digital Companions

Yuqian Wu, Zhijie Deng et al.

arXiv 2026 · 2026

LifeSide simulates lifelong digital companions using multi-session Memory Emotion Environment loops over 2,000 personas and 111K tasks to stress test cross session user modeling. LifeSide shows that even systems saturating existing memory benchmarks still fail to sustain accurate user understanding and emotional companionship over long horizons.

BenchmarkAgent MemoryLong-Term MemoryMemory Architecture

Lightweight LLM Agent Memory with Small Language Models

Jiaquan Zhang, Chaoning Zhang et al.

· 2026

LightMem orchestrates SLM-1 Controller, SLM-2 Selector, SLM-3 Writer, and STM MTM LTM stores to modularize retrieval, writing, and offline consolidation. On LoCoMo, LightMem reaches 34.50 F1 for GPT-4o multi hop questions, +1.64 over A-MEM, while keeping median retrieval latency at 83 ms.

Long-Term Memory

LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents

Siddharth Sharma, Nilesh Prasad Pandey et al.

arXiv 2026 · 2026

LIMBO combines a memory bank of successful interaction trajectories, feature extraction, and a two head LinUCB controller to choose replay mode and compute budget per task. On LifelongAgentBench DB and OS environments, LIMBO nearly matches fixed k=16 and LongLLMLingua baselines while reducing inference cost by up to ∼83% (∼53% on average).

Memory Architecture

LMEB: Long-horizon Memory Embedding Benchmark

Xinping Zhao, Xinshuo Hu et al.

· 2026

LMEB standardizes long-horizon memory retrieval using a unified IR-style format, four memory types, and a diverse dataset collection with components like LMEB-Episodic, LMEB-Dialogue, LMEB-Semantic, and LMEB-Procedural. LMEB’s headline result is that bge-multilingual-gemma2 with instructions achieves a Mean (Dataset) NDCG@10 of 61.41, while LMEB and MTEB retrieval scores have a Pearson correlation of -0.115, proving orthogonal capabilities.

Benchmark

Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents

Yifei Li, Weidong Guo et al.

· 2026

LoCoMo-Plus evaluates cognitive memory using an implicit Cue–Trigger Query Construction pipeline plus Semantic Filtering, Cue Memory Elicitation Validation, and insertion into LoCoMo dialogues. On LoCoMo-Plus, even strong systems like gemini-2.5-pro reach only 26.06% cognitive accuracy versus 71.78% factual accuracy on LoCoMo, exposing a large unresolved gap.

RAG

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

Di Wu, Zixiang Ji et al.

arXiv 2026 · 2026

LongMemEval-V2 evaluates long-term agent memory via context gathering over massive web-agent haystacks, using components like AgentRunbook-R, AgentRunbook-C, raw state slice pools, state transition event pools, and procedure and hint note pools. On LME-V2-Medium, AgentRunbook-C achieves 70.1% overall accuracy, a +24.2 point gain over the strongest RAG baseline with trajectory notes (45.9%).

Long-Term Memory

LPC-SM: Local Predictive Coding and Sparse Memory for Long-Context Language Modeling

Keqin Xie

· 2026

LPC-SM combines local attention, dual-timescale memory, predictive correction, Orthogonal Novelty Transport, and multi-head-coupled residual routing (mHC) inside a single autoregressive block. On OpenWebMath-10k continuation, LPC-SM with adaptive sparse control reaches final LM loss 10.787 versus 12.137 for a fixed sparse controller, a 12.517% improvement.

Benchmark

LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models

Xianglong Shi, Ruijie Yang et al.

arXiv 2026 · 2026

LSTMem equips each Transformer layer with cell memory C, hidden memory D, memory readout and attention correction, gated accumulation and expression, and cross-layer feedback to separate memory retention from expression and connect memory across depth. On HotpotQA with Qwen3-4B-Instruct, LSTMem reaches 68.12 F1 versus 60.47 for δ-Mem (MSW), and raises MemoryAgentBench average to 45.08 versus 38.85.

Agent Memory

LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation

Dongfang Li, Zixuan Liu et al.

arXiv 2026 · 2026

LycheeMemory V2 combines Online Semantic Segmentation, Segment-Level Memory Encoding, Structured Evidence Organization, and Plan-Guided Multi-Route Retrieval to batch coherent dialogue segments into typed, indexed records. On LoCoMo, LycheeMemory V2 achieves 89.22% overall accuracy versus 68.83% for A-Mem, and on LongMemEval-S it reaches 92.20% versus 71.60% for A-Mem, with up to 7.2× fewer construction tokens.

Agent Memory

MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents

Dongming Jiang, Yi Li et al.

· 2026

MAGMA organizes agent memory with an Intent-Aware Router, Adaptive Topological Retrieval, a Data Structure Layer of Relation Graphs and Vector Database, plus dual-stream Synaptic Ingestion and Asynchronous Consolidation. On LoCoMo, MAGMA achieves a 0.700 overall LLM-as-a-Judge score versus 0.590 for Nemori, and reaches 61.2% average accuracy on LongMemEval versus 56.2% for Nemori.