Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 12 of 19

Agent Memory

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents

Chidera Biringa, Lucas Yannul et al.

arXiv 2026 · 2026

Stashbird builds episodes, semantic relations, preference traces, community summaries, and persisted graph state into a single provenance-linked memory substrate for conversational agents. On LoCoMo with GPT-4.1-mini, Stashbird reaches 87.5% overall accuracy while using 44.6M vs 0.484M ingestion prompt tokens for Mem0 and 37M vs 0.484M for Graphiti, and 6.58M vs 0.484M for Hindsight.

Memory Architecture

StructMem: Structured Memory for Long-Horizon Behavior in LLMs

Buqiang Xu, Yijun Chen et al.

· 2026

StructMem organizes conversational history via Event-Level Binding, Cross-Event Consolidation, Dual-Perspective Extraction, and Temporal Anchoring to maintain temporally grounded relational events. On LoCoMo, StructMem achieves 76.82 overall accuracy while using just 1.937M construction tokens, improving over Mem0g’s 68.44 overall with 35.825M tokens.

Agent Memory

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Yu Cheng, Jiuan Zhou et al.

· 2026

TAME stores experiences as (query, experience, usefulness, trust annotation) in a shared Strategy Memory Bank, orchestrated by an Executor–Evaluator loop with feedback-driven memory evolution. On the GPT-5.2 AIME benchmark, TAME reaches 0.733 accuracy versus 0.587 for ReasoningBank, a +0.146 improvement while preserving trustworthiness on Trust-Memevo.

Agent Memory

TA-Mem: Tool-Augmented Autonomous Memory Retrieval for LLM in Long-Term Conversational QA

Mengwei Yuan, Jianan Liu et al.

· 2026

TA-Mem processes long conversations via an Episodic Memory Constructor, Multi-Indexed Database with Tools, and Memory Retrieval Agent that cooperate to chunk, index, and query structured memory pages. On the LoCoMo dataset, TA-Mem achieves 55.95 F1 and 51.47 BLEU-1 on temporal questions, beating Mem0 and MemoryOS while using 3755 tokens on average.

RAG

Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes

Neeraj Yadav

arXiv 2026 · 2026

MemStrata combines a deterministic assertion path, bi temporal ledger, production triple extractor, and strict_object_supersede gate to track current code facts across GitHub histories. On 130 SWE bench Lite plus Verified atomic transitions, MemStrata reaches 0.908–0.985 accuracy versus naive_rag’s 0.569–0.615 while reducing stale fact error from 0.361–0.377 to ≈0.

Long-Term Memory

TEPA: Revoking Stale Memories for Conflict-Robust Language Agents

Yan Zhou, Yue Ouyang et al.

arXiv 2026 · 2026

TEPA converts observations into keyed precedents, maintains an active set, a revoked archive, uses a retriever, and applies trial-validated promotion to control memory validity. On controlled hidden-regime drift, TEPA achieves 0.950 full-reversal success on the drift benchmark, a +0.740 improvement over last-write-wins at 0.210.

Agent Memory

The Compaction Cliff in Long-Running AI Agent Memory

Saber Zerhoudi, Jelena Mitrovic, Michael Granitzer

arXiv 2026 · 2026

Knowledge Triage combines a five-type classifier, TypeCompact, TypeDecompose, and TypeRetrieve with a SafetyMargin constraint detector to route agent memory through per-type retention policies. On AgentArtifactCorpus and SafetyMed, Knowledge Triage reaches 1.00 constraint preservation at 50% compaction and 97.0% pass rate, compared to 0.53 recall and 92.5% pass for Sonnet 4.6 /compact.

Long-Term Memory

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

Jundong Hu, Shekar Ramachandran

arXiv 2026 · 2026

The Memory Trust Gap evaluates persistent-memory agents using a Benefit/Safety suite, a trap sweep, a memory feature factorial, and a representation comparison across Qwen3 model sizes. It shows stale-value reliance stays as high as 1.00 while net harm ∆mem becomes strongly negative (down to −1.00) once agents are capable enough for the stale memory to override an authoritative tool.

Long-Term Memory

ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents

Cai Ke, Xin Liu et al.

arXiv 2026 · 2026

ThinkFlow compresses conversations into Probabilistic Latent Memory Skills, filters them via the Gated Latent Consolidator, and aligns them with the Context-Aware Hyper-Aligner under a Self-Supervised Test-Time Evolution paradigm. On the CC dataset with Qwen3-8B, ThinkFlow achieves 2.18 BLEU-4 and 77.20 Mauve, surpassing LD-Agent’s 1.27 BLEU-4 and 48.21 Mauve.

Agent Memory

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

Natchanon Pollertlam, Witchayut Kornsuwannawit

arXiv 2026 · 2026

Total Recall at What Cost instruments Mem0, Hindsight, and Mastra Observational Memory with a separable log–log cost model, synthetic dialogue grid, and matched LoCoMo accuracy. Total Recall at What Cost finds accuracy between 0.214–0.541 and serving-cost break-even points from immediate for Mastra OM to beyond 400 turns for Hindsight against the full-history baseline.

Long-Term Memory

Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents

Chuanchao Zang, Jianing Wang et al.

arXiv 2026 · 2026

PIPEPOISON combines stage-level signals, chain-structured losses, stability-calibrated configuration weights, and weighted stage losses to optimize poisoning content across diverse shadow pipelines. On 12 matched memory–agent configurations, PIPEPOISON reaches 73.4% Attack Utilization Rate on LongMemEval, LoCoMo, and BEAM, beating the best baseline (MemMorph) by 19.1 percentage points.

Benchmark

TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory

Tianyu Yang, Sudipta Paul et al.

arXiv 2026 · 2026

TRUSTMEM combines a Memory Transition Verifier, Memory Executor, Policy Model, Reference Model, and Retriever to score and optimize each memory transition for coverage, preservation, and faithfulness. TRUSTMEM achieves 65.7 average on MemoryAgentBench, improving Mem-α’s 59.2 by 6.5 points while reducing omission, corruption, and hallucination rates by up to 79.1%.

Benchmark

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

Arulnidhi Karunanidhi

arXiv 2026 · 2026

Aegis combines a four-stage content screening pipeline with provenance-weighted ranking over a memory store, using trust labels and semantic similarity to manage agent memory. Under a weak poisoning attack on LongMemEval_S, Aegis drops from 0.850 to 0.300 accuracy undefended, and even the shipped provenance weight (wt=0.15) is statistically indistinguishable from no defense (p=0.80).

Long-Term Memory

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

Peijun Qing, Fobo Shi, Soroush Vosoughi

arXiv 2026 · 2026

UTILMEM benchmarks long-term conversational memory utilization by combining Evidence Synthesis, Question Generation, Evidence-Sensitivity Filtering, and Haystack Construction over multi-session histories. UTILMEM’s main result shows NaiveRAG + Qwen3-Embedding-8B averages 59.3 Normalized Robustness with 49.2 Degradation Rate across five domains, while compression-heavy systems like Mem0 fall to 18.9 NR and 97.8 DR.

Agent Memory

VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents

Yuhao Chen, Yi Xu et al.

· 2026

VehicleMemBench builds multi-user interaction histories via Persona Group Generation, Event Chain Construction, Temporal Interleaving, and Conversation Generation, then evaluates agents in an executable in-vehicle environment. On VehicleMemBench, Gemini-3-Pro-Preview reaches 90.60 Exact State Match with Gold Memory but drops to 64.80 under Recursive Summarization, a 25.8-point loss that exposes memory as the main bottleneck.

Long-Term Memory

VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

Liyang Fan, Yingcheng Shi et al.

arXiv 2026 · 2026

VibeMemBench builds 111 SWE-style coding targets from 3,634 repository histories using the SIEVE pipeline, frozen verified experience, and a controlled memory intervention that only toggles the declared memory condition. On these targets, frozen verified experience raises Resolved by up to +4.5 percentage points for kimi-k2.7-code, while Mem0, SimpleMem, MemoryOS, and A-MEM fail to exceed the memory-off baseline in 11 of 12 solver–system pairings.

Agent Memory

Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents

Quang Dao, Purvi Kathalkar, Kenneth Eaton

arXiv 2026 · 2026

Weighted Memory Tree organizes agent history into a Hierarchical Memory Tree, Dynamic Retention Scoring, Memory Controller and Lifecycle Operations, and a Utility-Aware Prompt Synthesizer that jointly decide which memories stay active. On GAIA-Text, Weighted Memory Tree reaches 33.86% accuracy with Qwen3-8B versus 20.47% for Linear History, while reducing prompt-token usage from 43.67M to 32.48M.

Benchmark

What to Remember, What to Reveal: Privacy-Aware Memory for Conversational Agents

Wenjie Wang, Wenhe Si et al.

arXiv 2026 · 2026

SP-Mem uses Privacy-Aware Memory Writing, Partitioned Storage with Privacy Mapping, and Privacy-Aware Query-Time Reasoning and Authorized Retrieval over a hybrid vector plus graph memory layer to separate sanitized facts from exact private values. On the privacy-aware benchmark, SP-Mem reaches 0.996 privacy-entity identification accuracy and reduces unnecessary privacy usage on preference-only tasks from 16.00% for Full-context to 0.33%.

Long-Term Memory

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

Shweta Mishra, Shashank Mishra

arXiv 2026 · 2026

MERIT combines episodic tool-use tasks, memory conditions C0–C5, controlled corruption, and cost-adjusted metrics to stress-test long-term memory in agents. MERIT shows dependent-task success jumping from a leak-verified 0.00 floor to 0.55–1.00, while embedding retrieval drops to 0.30–0.95 on updated facts and update-on-write memories stay at 0.70–1.00.

Benchmark

WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems

Jiangnan Yu, Kisson Songqi Lin, Jilong Wu

arXiv 2026 · 2026

WhenLoss evaluates long-context memory systems by running a fixed reader under Truncated Full Context, Oracle Evidence, Complete Stored Memory, and Retrieved Memory to expose write and retrieval bottlenecks. On 500 LongMemEval questions, WhenLoss with Expected Predictive Compression reaches 0.49 CSM Contains Match vs 0.44 for Summary LLM, reducing the write-side gap to 0.04.

Agent Memory

When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory

Minkyu Song

arXiv 2026 · 2026

Dependency-aware Semantic Garbage Collection (DSGC) augments similarity-only retention with one-hop propagation over a supplied prerequisite graph E, using semantic relevance scores ri, propagation term πi, and a greedy retained subset S under budget. On the fixed benchmark with structurally indirect prerequisites, DSGC lifts full-chain retention from 0.03 to 0.90 under the lexical encoder and from 0.23 to 1.00 under the sentence encoder, compared to the similarity-only baseline.

Benchmark

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

Wen-Yu Chang, Yun-Nung Chen

arXiv 2026 · 2026

LOCOMO-CONV rewrites LoCoMo questions into dialog, implicit, counterfactual, and composed queries, then scores retrieval recall and response quality for systems like AnchorMem, A-MEM, mem0, Memora, and Naive RAG. On this benchmark, AnchorMem with multi-facet query rewriting achieves 0.754 dialog retrieval recall@10 on LoCoMo-Conv, a +0.095 improvement over vanilla AnchorMem, while abstractive systems like Memora reach 0.445–0.387 recall on implicit/composed queries.