Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 3 of 19

Benchmark

AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution

Yutao Yang, Junsong Li et al.

arXiv 2026 · 2026

AutoSkill composes prompt-driven modules Query Rewriting, Hybrid Skill Retrieval, Skill Extraction, Skill Management Decision, and Versioned Skill Merging to externalize user preferences and workflows into SKILL.md artifacts stored in a SkillBank. AutoSkill builds a multilingual SkillBank with 1,858 skills across four WildChat-1M subsets, showing that explicit skill artifacts can be continuously refined (e.g., professional_text_rewrite at version 0.1.34) without any parameter updates.

Memory Architecture

Auxiliary-predicted Compress Memory Model(ApCM Model): A Neural Memory Storage Model Based on Invertible Compression and Learnable Prediction

Weinuo Ou

· 2026

Auxiliary-predicted Compress Memory Model (ApCM Model) combines an Invertible Dimensionality Reduction and Predictor (IDRP) module with a Memory Read-Write Controller, including a global Memory Bank, cosine-similarity read, and access-frequency write policy. ApCM Model achieves lower MSE (0.987171 vs 1.001440) than a Key-Value Memory Network while compressing memory from 1024 to 128 dimensions on random data.

Benchmark

Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems

Soham Gadgil, David Alexander et al.

arXiv 2026 · 2026

Bad Memory evaluates how auto-loaded and referenced workspace files like CLAUDE.md, AGENTS.md, and behaviors.md affect agent behavior across sessions using a synthetic workspace, probe sessions, and adversarial goals. Bad Memory finds attack success rates as high as 100% and persistence rates up to 100% across Claude Code and Codex agents, revealing distinct failure modes in memory-based prompt injection.

Agent Memory

Belief Memory: Agent Memory Under Partial Observability

Junfeng Liao, Qizhou Wang et al.

arXiv 2026 · 2026

BeliefMem maintains an external belief-based memory bank with Add, Merge, and Belief-aware Retrieval over attribute-level hypotheses. On LoCoMo, BeliefMem reaches 42.38 F1 with GPT-4o-mini, beating Mem0’s 40.99 F1, and on ALFWorld it attains 59.88% success rate vs ReadAgent’s 54.03% (+5.85).

BenchmarkAgent MemoryLong-Term Memory

Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks

Zexue He, Yu Wang et al.

· 2026

MEMORYARENA orchestrates Memory-Agent-Environment Loops, Multi-Session Working Flow, Bundled Web Shopping, Group Travel Planning, and Progressive Web Search to stress-test how agents store and reuse information across sessions. MEMORYARENA’s main result is that agents with near-saturated scores on long-context benchmarks like LoCoMo still obtain Task Success Rates as low as 0.00–0.12 across its four environments.

BenchmarkBenchmarkLong-Term Memory

BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs

Sangyeon Yoon, Sunkyoung Kim et al.

· 2026

BenchPreS combines Contexts, User Profiles, Preference Attributes, Gold Labeling, and an LLM-as-Judge framework to test context-aware preference selectivity in persistent-memory LLMs. BenchPreS shows GPT-5.2 reaches 87.33% Appropriate Application Rate on BenchPreS while still having a 40.95% Misapplication Rate compared to Gemini 3 Pro’s 86.48% Misapplication Rate.

RAG

Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents

Donghua Cai, Yongheng Deng et al.

arXiv 2026 · 2026

Threader organizes raw dialogue using Incremental Topic-Coherent Segmentation, Multi-View Segment Representation, and Multi-Signal Evidence Retrieval to access long-horizon memory without LLM rewriting. On LongMemEval with GPT-5-mini, Threader achieves 90.60% overall accuracy and 93.59% Knowledge-Update accuracy, beating EverMemOS by 6.84 and 5.80 points respectively.

Agent Memory

Beyond Memory Majority: Latent-Source Reasoning for Multi-Agent Memory Arbitration

Chenchen Lin, Wenhao Yuan et al.

arXiv 2026 · 2026

CAMA combines Neuro-Symbolic Evidence Assignment, Effective Independent Evidence Estimation, Factor-Level Conflict Arbitration, and Active Independent-Evidence Recovery to decouple correlated memories into latent evidence slots and recover missing sources. On LongMemEval with DeepSeek V4 Flash, CAMA reaches 87.9 EM compared to MADAM RAG’s 86.8 EM while also cutting Replication Sensitivity from 15.3 to 7.8.

RAG

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

Han Zhang, Zihao Tang et al.

arXiv 2026 · 2026

RHELM builds 10 persona trajectories with Profile Generation, LOOP, External Sources, Dialogue Synthesis, and Question Curation to create 11,764 turns and 2,180 documents over 629 days. On the RHELM benchmark, Claude Opus 4.5 with RAG at k = 20 reaches 38.1 average score with external data, revealing substantial gaps in long term memory compared to simpler daily life benchmarks.

Memory Architecture

Beyond the Context Window: A Cost-Performance Analysis of Fact-Based Memory vs. Long-Context LLMs for Persistent Agents

Natchanon Pollertlam, Witchayut Kornsuwannawit

· 2026

Beyond the Context Window compares Conversation Segmentation, Fact Extraction, Embedding and Storage, and Retrieval Mechanism in a Mem0-based memory system against long-context GPT-5-mini. On LongMemEval, Beyond the Context Window finds LC GPT-5-mini reaches 82.40% accuracy, 33.4 percentage points above the memory system baseline.

Benchmark

Breaking the KV Cache Bottleneck: Fan Duality Model Achieves O(1) Decode Memory with Superior Associative Recall

Yasong Fan

· 2026

Fan Duality Model (FDM) uses the Fan Operator, Local-Global Cache, Freeze-Scan Training, and Holographic Reference Beam Decoding to separate wave-like compression from particle-like associative recall. On WikiText-103, Fan Duality Model (FDM) reaches 64.9 perplexity with Freeze-Scan and 62.79 with holographic decoding, while achieving 0.966 MQAR accuracy compared to Transformer at 0.606.

Long-Term Memory

CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion

Zheling Tan, Jin Gao, Dequan Wang

arXiv 2026 · 2026

CABLE augments long-term memory systems with Antecedent-Oriented Query Generation, Dual Retrieval, Overlap Subtraction, and Verification and Graph Update to build a sparse directed graph of complementary links. On MA-LongMemEval with A-MEM and Qwen3.5-27B, CABLE raises the mean LLM-judge score from 59.33% to 65.33%, a +6.00 percentage point gain over the A-MEM baseline.

RAG

Can Agent Memory Systems Track Evolving State?

Xinyi Fan, Miri Liu et al.

arXiv 2026 · 2026

StateMem represents conversations as structured StateStore entries built by a TurnEncoder, updated via deterministic Rechecker passes, and queried through a guided test-time recomputation prompt. On StateMemBench, StateMem reaches 0.363 gold rate on DeepSeek-V4-Flash, a +0.158 gain over the best retrieval baseline (Dense at 0.205) and +0.164 over long-context (0.149).

Long-Term Memory

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

S M Asif Hossain, Ruksat Khan Shayoni, Md Kishor Morol

arXiv 2026 · 2026

CAPTURE combines a preference-hypothesis extractor, authenticity gate, multi-timescale ledger, quarantine, and safety-bounded selector with causal audit to track latent user state over long horizons. On the D-PREFGUARD benchmark, CAPTURE achieves 71.5% win rate and 11.5% poisoning versus 69.3% and 15.9% for Sup. Transformer-∆t.

Agent Memory

CAST: Character-and-Scene Episodic Memory for Agents

Kexin Ma, Bojun Li et al.

· 2026

CAST builds views, scenes, character profiles, a semantic index, and an episodic index to organize dialogue into person-conditioned event structures. CAST reaches 62.02 F1 and 81.21 J on LOCOMO open questions, beating Zep by +13.24 F1 and +8.32 J and vanilla RAG by +25.15 F1 and +24.61 J.

Benchmark

Causal Episodic Memory for Feedback-Driven Agent Repair

Khang Nhat Hoang Vo, Tam Minh Chu et al.

arXiv 2026 · 2026

MERIT combines a Deterministic Error Classifier, Online Dual-Polarity Memory, Error-Typed Hybrid Retriever, and Type-Reliability-Aware Retrieval Variant to reuse oracle-verified corrections across queries. On Spider, MERIT reaches 69.79% execution accuracy versus 66.34% for Iterative repair using the same Qwen2.5-7B-Instruct backbone and initial predictions.

RAG

Chain-of-Memory: Lightweight Memory Construction with Dynamic Evolution for LLM Agents

Xiucheng Xu, Bingbing Xu et al.

· 2026

Chain-of-Memory (CoM) combines Memory Construction and Retrieval, Dynamic Memory Chain Evolution, State-Aware Gating Evolution, and Adaptive Path Truncation to turn flat retrieved turns into coherent reasoning chains. On LongMemEval with Qwen3-32B, Chain-of-Memory (CoM) achieves 76.40% total accuracy versus 66.00% for RAG (turn), while using only 8.8k tokens compared to 119.6k for Full-Context.

Memory Architecture

CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning

Yongshi Ye, Tian Lan et al.

arXiv 2026 · 2026

CHIME separates experience into a Hierarchical Memory Bank with a Plan Bank, Execution Bank, a Credit Attribution Gate, and Credit-Aware Memory Evolution to update memories only where credit belongs. On the DeepSeek-V4-Flash backbone, CHIME attains 39.01% average eval accuracy across τ²-Bench, VitaBench, BrowseComp-ZH, and BFCL-v4, beating A-MapReduce (35.33%) by +3.68 percentage points.

Agent Memory

ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory

Yongye Su, Wujiang Xu et al.

arXiv 2026 · 2026

ChronoMem wraps ADK’s LocalMemoryService, SQLite memory store, Version Index, and Reranker module into a semantic version-control layer that snapshots agent memory on every write and rolls back via natural-language queries. ChronoMem reaches 55.1% rollback-consistent QA accuracy on MemoryAgentBench with Qwen2.5-7B, compared to 35.5% for the RAG-only baseline.