Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 8 of 19

Benchmark

MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents

Boyu Yang, Jiazheng Sun et al.

arXiv 2026 · 2026

MeClear evaluates retrieved memories via Local Counterfactual Screening, Counterfactual Cooperative Attribution, and Verified Query-Scoped Clearance to decide which records to suppress per query. On ten LoCoMo-based long dialogue memory pools, MeClear attains 85.9% target recall and 82.3% binary task recovery, a +25.5 percentage point gain over Leave-One-Out.

Agent Memory

MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare

Yihao Wang, Haoran Xu et al.

arXiv 2026 · 2026

MedMemoryBench builds a four stage pipeline with Patient Profile Construction, Disease Progression Event Generation, Multi turn Sessions Simulation, and Memory Extraction and Query Construction to synthesize long horizon, clinically grounded interactions. On MedMemoryBench Efficient vs Mixed, methods like Letta drop from 51.21% to 41.55% average accuracy, quantifying how memory saturation harms personalized healthcare agents.

Benchmark

MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories

Lauri Lovén, Jaakko Sauvola et al.

arXiv 2026 · 2026

MELD merges distributed wiki brains using a merge decision procedure, Patch, status CRDT, and semantic publish subscribe over a typed-link memory substrate. On HotpotQA distractor, MELD’s decentralized merge achieves recall@5 = 0.630 vs 0.619 for a centralized store and 0.595 for naive union, with about 11% less live storage than union.

Agent Memory

Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents

Yiting Shen, Kun Li et al.

· 2026

Mem2ActBench integrates Heterogeneous Data Integration, Fact Extraction and Grouping, Memory Evolution Chain Construction, and Memory-anchored Q&A Construction to create long, interruption-heavy sessions with grounded tool calls. On Mem2ActBench, oracle retrieval reaches F1 53.8 while the best passive hybrid retriever at k=5 reaches only 30.7, revealing a 23.1 F1 gap in memory utilization.

Agent Memory

MemAdapter: Fast Alignment across Agent Memory Paradigms via Generative Subgraph Retrieval

Xin Zhang, Kailai Yang et al.

· 2026

MemAdapter combines a Generative Subgraph Retriever, Anchored Alignment Module, Target Alignment Module, Unified Memory Space, and Agent Model to turn diverse memory states into explicit evidence subgraphs. On NarrativeQA with Qwen2.5-7B, MemAdapter achieves F1 61.59 compared to 56.91 for MemoryLLM and 50.01 for Mem0, while completing cross-paradigm alignment in about 13 minutes with less than 5% of training compute.

Benchmark

MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents

Yongxian Wei, Yilin Zhao et al.

arXiv 2026 · 2026

MemAgent coordinates a Memory Agent, content aware routing architecture, training data synthesis pipeline, InsightGraph, and a short term memory provider to manage heterogeneous memory providers for a task agent. On GAIA, WebWalkerQA, and xBench DS, MemAgent reaches 75.7% average accuracy versus 65.7% for No Memory, a +10.0 percentage point gain.

Memory Architecture

MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents

Jiajun Dong, Yutao Hu et al.

arXiv 2026 · 2026

MemArbiter decomposes interaction histories into atomic memory items, organizes them into Memory Banks, and arbitrates their Dual-Band Memory Representation, Decision-Relevance Signals, and Temporal Presentation Gate before Prompt Assembly. On ALFWorld, MemArbiter achieves 92.54% SR@50 under a 750-token budget, beating Flat Retrieval by 25.38 percentage points.

Long-Term Memory

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

Jiadong Zhang, Xiaosong Ma

arXiv 2026 · 2026

MEMARENA builds an ego-centric, activity-dense benchmark using MASIM world simulation, PP sessions, PA sessions, ego-centric projection, and benchmark question generation to stress-test personal memory assistants. On MEMARENA-L, MEMARENA shows that at Qwen3-0.6B, switching from Memobase to MemSearch improves Recall from 23.7 to 56.2 (+32.5 pp) and Reasoning from 22.2 to 41.4 (+19.2 pp), while memory search adds only 87/7/48 ms overhead for BM25-RAG/Memobase/MemSearch on a Spark GB10 node.

Agent Memory

MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection

Zhewen Tan, Yilun Yao et al.

arXiv 2026 · 2026

MemAudit combines Counterfactual Memory Influence Score, Memory Consistency Graph, and a fused Detoxification Score to rank and remove suspicious memories in MINJA-poisoned agents. On MINJA QA with GPT-4o, MemAudit reduces attack success rate from 70.0% to 0.0%, beating random deletion, retrieval-frequency deletion, and nearest-neighbor contradiction filtering.

Agent Memory

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

Ruike Cao, Fanyu Zhao et al.

arXiv 2026 · 2026

MemCalib combines a 15,000-example benchmark, rubric-guided LLM-as-a-Judge evaluation, and the MemCalib-RL training algorithm built on ordered reward channels, bidirectional counterfactual localization, and channel-wise advantage redistribution. On the MemCalib test set, MemCalib-RL raises Sample Calibration Score to 81.12 on Qwen3.5-35B-A3B, a 1.73-point gain over GDPO while also improving Exact Calibration by 2.65 points.

Long-Term Memory

MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts

Zhen Tao, Jinxiang Zhao et al.

arXiv 2026 · 2026

MemConflict builds structured user profiles, simulates long-horizon timelines, injects conflicts, and evaluates systems via black-box Answer Accuracy and white-box Support Evidence Hit and Support Rank Score. On the MemConflict benchmark, MemOS attains 0.5539 average Answer Accuracy, 0.6710 SEH@3, and 0.5879 SRS, revealing large gaps versus systems like LangMem at 0.2822 AA.

BenchmarkBenchmarkBenchmarkAgent MemoryLong-Term Memory

MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents

Weiwei Xie, Shaoxiong Guo et al.

· 2026

MemEvoBench combines Misleading Memory Injection, Noisy Tool Returns, Biased User Feedback, and a Memory Modification Tool (+ModTool) to stress-test long-term memory safety in LLM agents across 7 domains and 36 risk types. On the QA Style benchmark, MemEvoBench shows Gemini-2.5-Pro’s ASR drops from 67.0% (Vanilla) to 19.0% with +ModTool in Round 1, while biased feedback can push GPT-5’s QA ASR from 59.0% to 78.0% by Round 3.

Agent Memory

Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory

Zhenting Wang, Huancheng Chen et al.

· 2026

Memex(RL) optimizes Indexed Experience Memory, CompressExperience, ReadExperience, and ContextStatus so Memex keeps only an indexed summary in-context while archiving full artifacts externally. On modified ALFWorld, Memex(RL) lifts task success from 24.22% to 85.61% over the Memex agent without RL while reducing peak working context from 16,934.46 to 9,634.47 tokens.

Agent Memory

MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing

Han Chen, Zining Zhang et al.

arXiv 2026 · 2026

MemForest combines parallel extraction, a shared memory substrate of canonical facts, scoped MemTrees, and forest recall plus tree browse to maintain temporal agent memory efficiently. On LongMemEval-S, MemForest with Qwen3-30B-A3B-Instruct-2507 achieves 81.8% pass@1 overall, 14.8 points above EverMemOS, while its input-normalized build rate is 6.0× higher.

Agent Memory

MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging

Junxi Wang, Te Sun et al.

arXiv 2026 · 2026

MemForest compresses agent memories by combining EventTree Semantic-Temporal Partitioning, EventTree Progressive Merging, and Anchor-Guided Propagation Retrieval into a unified EventTree-based memory forest. MemForest retains 97.1% of Mem0 performance and 99.7% of M3-Agent performance at 50% compression on LoCoMo, LongMemEval, PersonaMem, M3-Bench-robot, and M3-Bench-web while speeding up retrieval by up to 2.24× over Mem0 and M3-Agent.

Long-Term Memory

Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents

Yuanchen Bei, Tianxin Wei et al.

· 2026

Mem-Gallery organizes multi-session, multimodal conversations and tasks via Benchmark Construction, a unified Conversational Environment for Memory, and an Evaluation Framework spanning extraction, reasoning, and knowledge management. On Mem-Gallery, MuRAG reaches 0.6966 overall F1 and 0.8856 LLM-Judge on visual-centric search, surpassing NaiveRAG by 0.0992 overall F1.

Long-Term Memory

MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios

Yihang Ding, Wanke Xia et al.

· 2026

MemGround evaluates long-term memory via Surface State Memory, Temporal Associative Memory, and Reasoning-Based Memory in TRPG, No Case Should Remain Unsolved, and Type Help scenarios. MemGround’s main result shows frontier GPT-5.2 reaches 51.51% QA Overall in TRPG but only 23.61% QA Overall in Type Help, revealing severe reasoning degradation versus closed-source baselines like Gemini-3-Pro-Preview.

Long-Term Memory

MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models

Hyeonjeong Ha, Jeonghwan Kim et al.

arXiv 2026 · 2026

MEMGUARD decomposes conversations into type-specific memory atoms via Type-Aware Knowledge Decomposition, Self-Verified Extraction, Relational Knowledge Graph, and Query-Adaptive Type Routing stored in type-isolated memory. On HaluMem, MEMGUARD attains 89.53% anti-hallucination accuracy and 94.15 F1, improving anti-hallucination by 28.27 percentage points over MemOS.

Agent Memory

MemGym: a Long-Horizon Memory Environment for LLM Agents

Wujiang Xu, Yu Wang et al.

arXiv 2026 · 2026

MemGym wires environments like τ2-bench, SWE-Gym, WebArena-Infinity, MEMGYM-DR, and MEMGYM-CODEQA through a shared BaseMemoryManager, BaseAgent, BaseRunner, and MEMRM reward model. MemGym’s MEMRM achieves AUROC 0.985 on SWE-Gym compression events, turning multi-minute Docker rollouts into sub-second scalar memory-quality scores.

Agent Memory

MemLineage: Lineage-Guided Enforcement for LLM Agent Memory

Ciyan Ouyang, Rui Hou

arXiv 2026 · 2026

MemLineage wraps a single memory store with six modules: Provenance metadata, Ed25519 signing, an RFC 6962 Merkle log, a weighted lineage DAG, verifier-aware retrieval, and a sensitive-action gate that enforces Untrusted-Path Persistence. On a deterministic harness with AgentPoison-style, MemoryGraft-style, and sleeper-via-derivation attacks, MemLineage is the only configuration that achieves 0.00 ASR on all three while keeping per-operation overhead below one millisecond.