Paper archive

All AI memory papers

Browse 314 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 6 of 14

Memory Architecture

OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents

Yulin Hu, Zimo Long et al.

· 2026

OP-Bench benchmarks over-personalization using Irrelevance, Repetition, and Sycophancy across 1,700 long-horizon queries built from LoCoMo, then analyzes memory systems like RAG, LDAgent, Mem0, MemU, and MEMOS. Self-ReCheck, a lightweight memory filter, improves average OP-Bench scores by 29% for Qwen3-8B while preserving personalization on LoCoMo compared to BASE.

Long-Term Memory

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

Bowen Yang, Kaiming Jin et al.

arXiv 2026 · 2026

OS-SYMPHONY coordinates an Orchestrator, Reflection-Memory Agent, and Versatile Tool Agents (Multimodal Searcher, Grounders, Coder) to stabilize long-horizon GUI workflows and fetch visual tutorials on demand. On OSWorld-Verified, OS-SYMPHONY with GPT-5 scores 65.84% at 100 steps, beating Agent S3 w/ GPT-5 (62.63%) by 3.21 percentage points.

Long-Term Memory

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory

Zhifei Xie, Zongzheng Hu et al.

· 2026

Pask combines Demand Detection (IntentFlow), Pask-MM hierarchical memory, and the Pask-PAS proactive agent system to continuously infer latent needs and act through tools and frontier models. On LatentNeeds-Bench, Pask’s IntentFlow achieves 84.2% balanced accuracy versus 80.8% for Gemini-3-Flash, under ~1.3–1.5 s per-turn latency.

Long-Term Memory

PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments

Shuochen Liu, Junyi Zhu et al.

· 2026

PERMA constructs temporally ordered Interaction Events, a dynamic Persona State, and a dual-process Memory System over Clean, In-session Noise, and Style-aligned Long-context datasets. PERMA shows that structured memory agents like MemOS can compress raw 34k-token histories down to ~700 tokens while preserving higher MCQ accuracy than vanilla RAG in realistic, noisy multi-session personalization tasks.

Long-Term Memory

PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?

Sidharth Pulipaka, Oliver Chen et al.

· 2026

PersistBench combines Monte Carlo Tree Search, Seed Initialization and Candidate Generation, Search and Scoring, and Human Verification to build realistic long-term memory test cases across cross-domain leakage, sycophancy, and beneficial memory use. PersistBench then reports a median 53% cross-domain leakage failure rate and 97.8% sycophancy failure rate on 18 LLMs using its 500-sample benchmark.

Long-Term Memory

PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents

Ke Yang, Zixi Chen et al.

· 2026

PLUGMEM standardizes episodic traces with the Structuring Module, retrieves via an abstraction-aware Retrieval Module, and compresses outputs through a Reasoning Module into a unified memory graph. On HotpotQA, PLUGMEM reaches 61.4 EM and 74.1 F1 with only 81.6 memory tokens, compared to 51.7 EM and 62.7 F1 with 659.2 tokens for Vanilla Retrieval.

Agent Memory

Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents

Wei Zou, Mingwen Dong et al.

· 2026

eTAMP stores attacker-crafted payloads inside Trajectory Memory, then reactivates them via semantic Cross-Site Task Pairing and Chaos Monkey induced frustration during later tasks. On (Visual)WebArena, eTAMP achieves up to 32.5% ASRB on GPT-5-mini and 23.4% on GPT-5.2 under Frustration Exploitation, showing persistent cross-session compromise without direct memory access.

Benchmark

QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents

Heng Wang, Yifei Li et al.

arXiv 2026 · 2026

QUMem combines Dynamic Episode Construction, Typed Memory Decomposition, and Query-Conditioned User-State Inference to segment histories into semantically coherent episodes and split them into factual, preference, and transferable insight memories. On PersonaMem, QUMem reaches 70.58% overall accuracy with Gemini-3.5-flash, improving over Mem0’s 63.29% and Zep’s 54.46%.

RAG

RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation

Kyle Wild, Yusuke Takahashi, Asako Uraki

arXiv 2026 · 2026

Ingest-Time Semantic Compilation (ISC) builds a semantic substrate in PostgreSQL using a geometric layer, symbolic layer, validation gate, and index_outbox for incremental maintenance and migration. On 500 broadcast-interview transcripts, ISC’s compiled claims reach 85.2% accuracy from roughly 2.2k reader tokens, beating the best chunk configuration at 72.5% from roughly 16.3k tokens and matching a contextualized stack that spends ~47.7k tokens.

Long-Term Memory

REAL: A Reasoning-Enhanced Graph Framework for Long-Term Memory Management of LLMs

Keer Lu, Liwei Chen et al.

arXiv 2026 · 2026

REAL represents conversational memory as a temporal directed property graph using Atomic Fact Extraction, Incremental Graph Update with Non-Destructive Temporal Evolution, Confidence Stratification and Upgrade, and Exploration Intent Enrichment. On LoCoMo, REAL with DeepSeek-V3 reaches 59.98% EM versus 55.12% EM for A-MEM, and averages a 22.72% improvement over existing memory management methods.

Agent Memory

RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction

Haonan Bian, Zhiyuan Yao et al.

arXiv 2026 · 2026

RealMem constructs realistic long-term project dialogues via Project Foundation Construction, Multi-Agent Dialogue Generation, and Memory and Schedule Management over eleven scenarios and 2,000+ cross-session dialogues. On RealMem, the Oracle QA Score reaches 0.804 while the strongest memory system, MemoryOS, achieves only 0.567, quantifying the difficulty of real-world project memory.

Long-Term Memory

Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents

Baichuan Li, Junyi Yao, Zihao Zheng

arXiv 2026 · 2026

Memory–Clarification Boundary (MCB) uses persist, ephemeral, verify, and clarify actions plus the MCB-Act tool-call variant to probe memory commitment decisions in LLM agents. On the 70-item held-out MCB test, Qwen3.5-9B few-shot prompting raises accuracy from 0.557 to 0.771 while Qwen3.5-9B act-mode accuracy falls to 0.343, showing that stated decisions and tool-call choices diverge.

Long-Term Memory

REMem: Reasoning with Episodic Memory in Language Agent

Yiheng Shu, Saisri Padmaja Jonnalagedda et al.

· 2026

REMem converts interaction histories into a hybrid memory graph via Gist Extraction, Fact Extraction, and Graph Construction, then queries it with Agentic Inference tools like semantic retrieve and find entity contexts. On the Test of Time benchmark, REMem-I achieves 93.1% EM compared to 66.9% for HippoRAG 2, a +26.2 point gain in episodic reasoning.

Agent Memory

Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey

Wei-Chieh Huang, Weizhi Zhang et al.

arXiv 2026 · 2026

Rethinking Memory Mechanisms of Foundation Agents in the Second Half organizes foundation agent memory using a three-dimensional taxonomy of Memory Substrates, Memory Cognitive Mechanisms, and Memory Subjects plus an operation and optimization view. Rethinking Memory Mechanisms of Foundation Agents in the Second Half synthesizes 218 memory-related agent papers from 2023 Q1–2025 Q4, highlighting the sharp acceleration of external memory and working or episodic mechanisms in 2025.

Benchmark

RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory

Jingbo Ji, Lingyi Li et al.

arXiv 2026 · 2026

RippleMem converts dialogues into Cue-Rich Episodic Memory Construction, organizes them in an Event-Centric Memory Graph, and queries them via Adaptive Associative Recollection and Evidence Assembly. On LoCoMo, RippleMem attains 87.14% LLM-as-a-Judge accuracy, 52.49% F1, and 44.05 BLEU-1, beating RF-Mem’s 83.83% judge accuracy.

Benchmark

STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?

Hanxiang Chao, Yihan Bai et al.

arXiv 2026 · 2026

STALE evaluates latent user-state tracking by probing State Resolution, Premise Resistance, and Implicit Policy Adaptation across 400 conflict scenarios packaged into long user-assistant histories. CUPMEM applies structured state consolidation and propagation-aware search on STALE, reaching 68.0% overall accuracy versus 55.2% for Gemini-3.1-pro.

Memory Architecture

StructMem: Structured Memory for Long-Horizon Behavior in LLMs

Buqiang Xu, Yijun Chen et al.

· 2026

StructMem organizes conversational history via Event-Level Binding, Cross-Event Consolidation, Dual-Perspective Extraction, and Temporal Anchoring to maintain temporally grounded relational events. On LoCoMo, StructMem achieves 76.82 overall accuracy while using just 1.937M construction tokens, improving over Mem0g’s 68.44 overall with 35.825M tokens.

Agent Memory

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Yu Cheng, Jiuan Zhou et al.

· 2026

TAME stores experiences as (query, experience, usefulness, trust annotation) in a shared Strategy Memory Bank, orchestrated by an Executor–Evaluator loop with feedback-driven memory evolution. On the GPT-5.2 AIME benchmark, TAME reaches 0.733 accuracy versus 0.587 for ReasoningBank, a +0.146 improvement while preserving trustworthiness on Trust-Memevo.

Agent Memory

TA-Mem: Tool-Augmented Autonomous Memory Retrieval for LLM in Long-Term Conversational QA

Mengwei Yuan, Jianan Liu et al.

· 2026

TA-Mem processes long conversations via an Episodic Memory Constructor, Multi-Indexed Database with Tools, and Memory Retrieval Agent that cooperate to chunk, index, and query structured memory pages. On the LoCoMo dataset, TA-Mem achieves 55.95 F1 and 51.47 BLEU-1 on temporal questions, beating Mem0 and MemoryOS while using 3755 tokens on average.

RAG

Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes

Neeraj Yadav

arXiv 2026 · 2026

MemStrata combines a deterministic assertion path, bi temporal ledger, production triple extractor, and strict_object_supersede gate to track current code facts across GitHub histories. On 130 SWE bench Lite plus Verified atomic transitions, MemStrata reaches 0.908–0.985 accuracy versus naive_rag’s 0.569–0.615 while reducing stale fact error from 0.361–0.377 to ≈0.