Paper archive

All AI memory papers

Browse 453 explainers on agent memory, RAG, long-term context, personalization, benchmarks, and memory architectures.

Page 16 of 19

Benchmark

SGMem: Sentence Graph Memory for Long-Term Conversational Agents

Yaxiong Wu, Yongyue Zhang et al.

· 2025

SGMem organizes long conversations via SGMem Construction and Management, SGMem Usage, sentence level graphs, and multi hop retrieval over sessions, rounds, turns, summaries, facts, and insights. SGMem achieves 0.700 Accuracy (Top 5) on LongMemEval and 0.526 on LoCoMo, beating the RAG-SMFI baseline at 0.676 and 0.510 respectively.

Long-Term Memory

StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns

Luanbo Wan, Weizhi Ma

· 2025

StoryBench evaluates long-term memory by embedding LLMs in Dynamic Narrative and Multi-Turn Decision-Making, two Task Modes for Evaluating LTM, and Tailored Metrics for Assessing LTM Models over a 311-scene, 86-choice graph. StoryBench shows Doubao1.5-pro achieving 80.98% Overall Acc in Immediate Feedback mode on The Invisible Guardian while Claude 3.5 Sonnet attains the highest Success Count of 8, revealing distinct strengths and weaknesses.

RAGMemory Architecture

TeleMem: Building Long-Term and Multimodal Memory for Agentic AI

Chunliang Chen, Ming Guan et al.

· 2025

TeleMem converts interactions into unified semantic nodes via the representation layer, organizes them in a memory graph with Insert and ReInsert, and reads them using closure-based retrieval and a ReAct-style multimodal agent. On ZH-4O, TeleMem reaches 86.33% QA Accuracy, beating the Mem0 baseline at 70.20% and the RAG baseline at 62.45%.

Memory Architecture

Test-time regression: a unifying framework for designing sequence models with associative memory

Ke Alexander Wang, Jiaxin Shi, Emily B. Fox

· 2025

Test-time regression uses memorization as regression, memory retrieval, and test-time regression layers to reinterpret sequence architectures as solving a regression problem over key value pairs during the forward pass. This unification shows how linear attention, state space models, fast weight programmers, online learning layers, and softmax attention are all instances of the same framework and explains phenomena like linear attention’s failures and the role of query key normalization.

PickMemory Architecture

Titans: Learning to Memorize at Test Time

Ali Behrouz, Peilin Zhong, Vahab Mirrokni

arXiv 2025 · 2025

Titans combines a Core short-term attention block, a deep Long-term Memory module, and Persistent Memory tokens, with three integration variants: Memory as a Context (MAC), Memory as a Gate (MAG), and Memory as a Layer (MAL). On language modeling and reasoning benchmarks, Titans (MAC) at 760M parameters achieves 52.51 average accuracy vs 51.49 for Gated DeltaNet-H2, while also solving BABILong tasks that defeat GPT-4.

Benchmark

TokMem: One-Token Procedural Memory for Large Language Models

Zijun Wu, Yongchang Hao, Lili Mou

· 2025

TokMem adds a Memory Bank of trainable memory tokens to a frozen Transformer backbone, using memory routing, conditional generation, and renormalization to store and recall procedures. On Super-Natural Instructions and APIGen function-calling, TokMem reaches 67.0 ROUGE-L and 99.1 tool-selection F1, surpassing Replay Memory and LoRA fine-tuning with far fewer parameters.

Benchmark

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

Amir Zandieh, Majid Daliri et al.

· 2025

TurboQuant combines MSE Optimal TurboQuant, Inner-product Optimal TurboQuant, QJL, Random Rotation Matrix Π, and Lloyd-Max Quantizer to quantize vectors online with near-optimal distortion-rate guarantees. TurboQuant matches the Shannon lower bound within a factor of √(3π/2)≈2.7 for MSE and achieves absolute quality neutrality for KV cache quantization at 3.5 bits per channel compared to full-precision baselines.

Benchmark

Unable to Forget: Proactive Interference Reveals Working Memory Limits in LLMs Beyond Context Length

Chupei Wang, Jiaqiu Vince Sun

· 2025

PI-LLM uses a synthetic key–value retrieval task, Interference Endurance Score (IES), per-key forget prompts, and a mock QA reset to stress-test working-memory-like behavior under proactive interference. PI-LLM finds a universal log-linear decay in retrieval accuracy across 0.6B–637B-parameter LLMs as interference grows, revealing that parameter size, not context window length, predicts interference robustness.

Memory Architecture

Understanding Transformer from the Perspective of Associative Memory

Shu Zhong, Mingyu Xu et al.

· 2025

Understanding Transformer from the Perspective of Associative Memory reframes Softmax Attention, Linear Attention, FFN, and DeltaNet as instances of a unified associative memory with explicit memory capacity and update rules. Using this lens, Understanding Transformer from the Perspective of Associative Memory derives retrieval SNR for different kernels, unifies attention and FFNs, and proves that DeltaFormer achieves circuit complexity beyond TC0, reaching NC1 expressivity.

RAG

Understanding Users' Privacy Perceptions Towards LLM's RAG-based Memory

Shuning Zhang, Rongjun Ma et al.

· 2025

Understanding Users' Privacy Perceptions Towards LLM's RAG-based Memory analyzes users' mental models, privacy calculus, and expectations around RAG-based memory across generation, management, usage, and updating. Understanding Users' Privacy Perceptions Towards LLM's RAG-based Memory finds users demand explicit consent, fine-grained editing and deletion, and visibility into inferred information to trust RAG-based memory systems.

Agent Memory

Unveiling Privacy Risks in LLM Agent Memory

Bo Wang, Weiyi He et al.

· 2025

MEXTRA crafts black box attacking prompts and automated diverse prompt generators that target the memory module, similarity scoring function, retrieval depth, memory size, and LLM backbone. MEXTRA extracts 50 queries from a 200 record EHRAgent memory and 26 from RAP, with extracted efficiency up to 0.42 compared to weaker baselines without workflow aligned prompts.

Agent MemoryMemory Architecture

WebATLAS: An LLM Agent with Experience-Driven Memory and Action Simulation

Jiali Cheng, Anjishnu Kumar et al.

· 2025

WebATLAS combines a Planner, Actor, Critic, and Multi-layered Memory (Working Memory, Cognitive Map, Semantic Memory) to simulate and score actions before executing them on the web. On WebArena-Lite, WebATLAS achieves 63.0% average success versus 53.9% for Plan-and-Act, a +9.1 point gain without website-specific fine-tuning.

BenchmarkBenchmarkMemory Architecture

WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning

Woongyeong Yeo, Kangsan Kim et al.

· 2025

WorldMM dynamically coordinates Episodic Memory, Semantic Memory, Visual Memory, an Adaptive Retrieval Agent, and a Response Agent to answer queries over hour- to week-long videos. On five long video QA benchmarks, WorldMM-GPT reaches 69.5% average accuracy, beating M3-Agent’s 55.1% by 14.4 points and the best prior memory baseline HippoRAG’s 57.0% by 12.5 points.

Benchmark

Zep: A Temporal Knowledge Graph Architecture for Agent Memory

Preston Rasmussen, Pavlo Paliychuk et al.

· 2025

Zep builds memory using Episode Subgraph, Semantic Entity Subgraph, Community Subgraph, and a three-stage Search–Reranker–Constructor pipeline over the Graphiti temporal knowledge graph. On LongMemEval, Zep with gpt-4o scores 71.2% vs a 60.2% full-context baseline and reduces average latency from 28.9 s to 2.58 s.

Benchmark

AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents

Petr Anokhin, Nikita Semenov et al.

· 2024

AriGraph builds a joint world model by combining semantic memory, episodic memory, semantic search, episodic search, and the Ariadne cognitive architecture into a single evolving graph. On NetHack, AriGraph lets Ariadne reach a score of 593.00 with room-only observations, compared to 341.67 for NetPlay with the same input and 675.33 for NetPlay with full level observations.

Memory Architecture

CAMELoT: Towards Large Language Models with Training-Free Consolidated Associative Memory

Zexue He, Leonid Karlinsky et al.

arXiv 2024 · 2024

CAMELoT augments frozen LLaMA2‑7B with Read, Augment, and Write associative memory modules that maintain consolidated key–value slots with novelty and recency‑aware updates. On Arxiv long‑context language modeling, CAMELoT reaches 3.60 average perplexity versus 5.12 for LLaMa2‑7B, a 29.7% reduction without any retraining.

RAG

Deciphering the Interplay of Parametric and Non-parametric Memory in Retrieval-augmented Language Models

Mehrdad Farahani, Richard Johansson

· 2024

Deciphering the Interplay of Parametric and Non-parametric Memory instruments causal mediation analysis, Experiment 1, Experiment 2, and Path Specific Effects (PSE) inside ATLAS to trace how parametric and non-parametric memories compete token-by-token. Deciphering the Interplay of Parametric and Non-parametric Memory reports a strong shift toward counterfactual answers in altered contexts, with a t-test p-value of 1.60e-4 and Cohen’s d of -0.9851 for non-parametric versus parametric behavior.

Memory Architecture

Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformers

Yibo Jiang, Goutham Rajendran et al.

· 2024

Do LLMs dream of elephants studies how a self-attention layer, value matrix, embedding matrix, latent concept association task, and context hijacking prompts interact to implement associative memory in transformers. Do LLMs dream of elephants proves theoretically (Theorem 1, Theorem 4) and empirically that a one-layer transformer can achieve arbitrarily small error on latent concept association by using the value matrix as associative memory.

Memory Architecture

Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models

Yabin Zhang, Wenjie Zhu et al.

arXiv 2024 · 2024

Dual Memory Networks combines a Dynamic Memory Network, Static Memory Network, a shared ReadOut module, Projection Layers ω, and a Memory Interactive Strategy to build sample-adaptive classifiers on top of frozen CLIP encoders. On zero-shot ImageNet with ViT-B/16, Dual Memory Networks achieves 72.25% accuracy vs 66.73% for CLIP and 68.98% for TPT, a +5.52 and +3.27 point gain respectively.

Benchmark

DynaMem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation

Peiqi Liu, Zhanqiu Guo et al.

· 2024

DynaMem maintains a Dynamic 3D Voxel Map, supports Embedded Vision Language Features and Multimodal Large Language Models querying, and exposes Exploration Primitives and an Obstacle map for navigation and manipulation. On real Stretch SE3 experiments, DynaMem achieves a 70% success rate on dynamic pick-and-drop tasks compared to 30% for the static OK-Robot baseline.

Benchmark

Efficient Episodic Memory Utilization of Cooperative Multi-Agent Reinforcement Learning

Hyungho Na, Yunkyeong Seo, Il-chul Moon

· 2024

EMU combines a semantic memory embedding via deterministic conditional autoencoder and an episodic incentive built from desirability in the episodic buffer to guide cooperative MARL. EMU improves learning speed and final win-rates on StarCraft II SMAC and Google Research Football compared to QPLEX, CDS, and EMC, especially on super hard maps.

RAG

Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation

Quanting Xie, So Yeon Min et al.

· 2024

Embodied-RAG builds a multimodal Topological Map and a hierarchical Semantic Forest and then runs Top-down Retrieval with LLM-based selection and hybrid re-ranking to drive Generation of waypoints and explanations. On the E-multimodal Embodied-Experiences dataset, Embodied-RAG reaches P(Q|A)=0.67 for implicit queries (Q only), compared to 0.13 for LightRAG, while building graph memory 9.76× faster than LightRAG.

Survey

Evaluating Very Long-Term Conversational Memory of LLM Agents

Adyasha Maharana, Dong-Ho Lee et al.

· 2024

LOCOMO builds very long-term multi-modal dialogues by combining Persona, Temporal Event Graph, Virtual Agent Architecture, Image Sharing and Image Reaction, and Human Verification and Editing. On the LOCOMO QA benchmark, GPT-3.5-turbo-16k with 16K context achieves 37.8 F1 overall, 50.9 F1 with RAG over dialogs, while humans reach 87.9 F1.