Category

Benchmark

Empirical studies and benchmarks on context, recall, and memory limitations in LLMs.

28 papers

Benchmark

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Xiangchen Cheng, Yunwei Jiang et al.

arXiv 2026 · 2026

AgenticSTS composes each decision from L1 protocol instructions, L2 state-typed prompts, L3 game knowledge, L4 episodic memory, and L5 skill library instead of an accumulating transcript. On fixed A0 Slay the Spire 2 runs, AgenticSTS with L5 skills wins 6/10 games versus 3/10 for the no-scaffold baseline, and climbs to A6–A8 in auto-mode streams.

Benchmark

AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment

Jianfei Xiao, Xiang Yu et al.

· 2026

AlpsBench combines Personalized Information Extraction, Personalized Information Update, Personalized Information Retrieval, and Personalized Information Utilization over 2,500 WildChat dialogues with human-verified structured memories. AlpsBench shows, for example, that Gemini-3 Flash scores 51.67 on Task 1 Extraction while DeepSeek Reasoner reaches 0.9569 retrieval recall with 100 distractors on AlpsBench.

BenchmarkBenchmarkLong-Term Memory

A-MBER: Affective Memory Benchmark for Emotion Recognition

Deliang Wen, Ke Sun, Yu Wang

· 2026

A-MBER builds multi-session conversational scenarios via a staged pipeline of persona specification, long-horizon planning, conversation generation, annotation, question construction, and benchmark-unit packaging. On A-MBER, a structured memory system reaches 0.69 judgment accuracy, 0.66 retrieval, and 0.65 explanation versus 0.34, 0.29, and 0.31 for a no-memory baseline.

BenchmarkAgent Memory

AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations

Cheng Jiayang, Dongyu Ru et al.

· 2026

AMemGym combines Structured Data Generation, On-Policy Interaction, Evaluation Metrics, and Meta-Evaluation to script user state trajectories, drive LLM-simulated role-play, and score write–read–utilization behavior. On AMemGym’s base configuration, AWE-(2,4,30) reaches a 0.291 normalized memory score on interactive evaluation, while native gpt-4.1-mini only achieves 0.203, exposing substantial gaps between memory agents and plain long-context LLMs.

BenchmarkAgent Memory

ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks

Samuel Sameer Tanguturi

· 2026

ATANT v1.1 structurally analyzes seven benchmarks using the 7 v1.0 continuity properties, the 10 checkpoints, a property-coverage matrix, and the Kenotic v1.0 reference implementation. ATANT v1.1 reports 96% ATANT cumulative-scale versus 8.8% LOCOMO substring accuracy, showing that LOCOMO, LongMemEval, BEAM, MemoryBench, Zep eval, MemGPT/Letta, and RULER measure different properties from continuity.

BenchmarkAgent MemoryLong-Term Memory

Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks

Zexue He, Yu Wang et al.

· 2026

MEMORYARENA orchestrates Memory-Agent-Environment Loops, Multi-Session Working Flow, Bundled Web Shopping, Group Travel Planning, and Progressive Web Search to stress-test how agents store and reuse information across sessions. MEMORYARENA’s main result is that agents with near-saturated scores on long-context benchmarks like LoCoMo still obtain Task Success Rates as low as 0.00–0.12 across its four environments.

BenchmarkBenchmarkLong-Term Memory

BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs

Sangyeon Yoon, Sunkyoung Kim et al.

· 2026

BenchPreS combines Contexts, User Profiles, Preference Attributes, Gold Labeling, and an LLM-as-Judge framework to test context-aware preference selectivity in persistent-memory LLMs. BenchPreS shows GPT-5.2 reaches 87.33% Appropriate Application Rate on BenchPreS while still having a 40.95% Misapplication Rate compared to Gemini 3 Pro’s 86.48% Misapplication Rate.

Benchmark

ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents

Yating Wu, Yuhao Zhang et al.

· 2026

ContextWeaver organizes agent histories using Dependency-Aware Context Construction, Dependency Summarizer, and a Validation and Test Layer into a dependency graph that drives selective context weaving. On SWE-Bench Verified, ContextWeaver with Claude Sonnet 4 achieves 66.0% pass@1 versus 63.2% for the Sliding Window baseline under the same window size.

Benchmark

DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings

Wenya Xie, Shengming Zhou et al.

arXiv 2026 · 2026

DynamicMem builds 15‑month, multi‑app trajectories using Multi-timescale user profile construction, Intent-conditioned event chain generation, and State-consistent multi-app log generation to stress-test long-horizon memory. DynamicMem’s main result shows State Completion declines by up to 26.5 points for A-Mem from checkpoint C1 to C5, while Personalized Service scores remain stable or improve, revealing hidden failure modes in memory systems.

Benchmark

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

Yijun Chen, Yaqi Zheng et al.

arXiv 2026 · 2026

EM2Mem organizes long videos into Event-Centric Multimodal Memory Cells, Temporal Context Views, and event-linked Episodic and Semantic Graphs for align-then-retrieve reasoning. On Video-MME (L), EM2Mem reaches 76.8% average accuracy, beating the strongest memory baseline WorldMM† at 73.1% and cutting total inference tokens by 63.66%.

Benchmark

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

Zhe Ren, Yibo Yang et al.

arXiv 2026 · 2026

GateMem structures shared-memory evaluation around Episodes and Memory State, Checkpoints and Governance Categories, and a multiplicative Memory Governance Score that couples utility, access violations, and forgetting failures. GateMem shows, for example, GPT-5.4 with LONG-CONTEXT reaches 80.1% MGS on the Medical domain while RAG-NAIVE only achieves 44.7%, yet even LONG-CONTEXT still leaks unauthorized or deleted information.

Benchmark

HippoCamp: Benchmarking Contextual Agents on Personal Computers

Zhe Yang, Shulin Tian et al.

arXiv 2026 · 2026

HippoCamp evaluates contextual agents on realistic personal file systems using structured trajectories with search, perception, and reasoning capability tags over 42.4 GB of multimodal data. On HippoCamp, ChatGPT Agent Mode reaches 48.3% profiling accuracy and 62.8% factual retention accuracy, while RAG and search agents lag far behind.

Benchmark

Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents

Yifei Li, Weidong Guo et al.

· 2026

LoCoMo-Plus evaluates cognitive memory using an implicit Cue–Trigger Query Construction pipeline plus Semantic Filtering, Cue Memory Elicitation Validation, and insertion into LoCoMo dialogues. On LoCoMo-Plus, even strong systems like gemini-2.5-pro reach only 26.06% cognitive accuracy versus 71.78% factual accuracy on LoCoMo, exposing a large unresolved gap.

BenchmarkBenchmarkBenchmarkAgent MemoryLong-Term Memory

MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents

Weiwei Xie, Shaoxiong Guo et al.

· 2026

MemEvoBench combines Misleading Memory Injection, Noisy Tool Returns, Biased User Feedback, and a Memory Modification Tool (+ModTool) to stress-test long-term memory safety in LLM agents across 7 domains and 36 risk types. On the QA Style benchmark, MemEvoBench shows Gemini-2.5-Pro’s ASR drops from 67.0% (Vanilla) to 19.0% with +ModTool in Round 1, while biased feedback can push GPT-5’s QA ASR from 59.0% to 78.0% by Round 3.

BenchmarkBenchmarkAgent Memory

MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization

Weizhi Zhang, Xiaokai Wei et al.

· 2026

MEMORYCD builds a user memory pool Mu from lifelong Amazon Review histories and evaluates long-context prompting, Mem0, LoCoMo, ReadAgent, MemoryBank, and A-Mem across rating, ranking, and personalized text tasks. On Books and Home & Kitchen, MEMORYCD shows GPT-5 reaches RMSE 0.551–0.624 and NDCG@3 up to 0.610, while Gemini-2.5 Pro peaks at ROUGE-L 0.222 for generation, revealing substantial remaining gaps to real user behavior.

Benchmark

Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning

Zhen Bi, Xueshu Chen et al.

arXiv 2026 · 2026

Memory Boundary-Aware Router combines Behavioral Scientific Knowledge Boundary Characterization, Knowledge-Circuit View of Internal Knowledge Boundaries, External Boundary-Aware Data Router, and Internal Boundary-Aware Parameter Router to control conditional memory inside Transformer blocks. On BioProBench Error Correction, Memory Boundary-Aware Router reaches 0.63 accuracy with Qwen3-8B, surpassing the 0.62 accuracy of Qwen3-8B-Memory-LoRA.

Benchmark

MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

Arseniy Varlamov, Rishat Zinnatullin et al.

arXiv 2026 · 2026

MemToC builds a controlled arbitration pipeline from ToolHop pool construction and filtering, Distractor generation, Correctness-conditioned evaluation, and Annotator quality control to separate source correctness from source preference. On MemToC, cross-fitted SFT lifts Llama-3.1-8B-Instruct correct-answer retention from 17.1% to 31.6% while keeping correct-tool following at 92.8–93.1%, and DPO achieves 22.3% retention with 92.1% correct-tool following.

Benchmark

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Mengru Wang, Haozhe Luo et al.

arXiv 2026 · 2026

MemTrapBench combines a Taxonomy, Instance Construction, and Two-Gate Quality Flow to generate 1,050 adversarial multi-turn dialogues that trigger Reasoning Fixation and Belief Distortion in LLM memory use. On Gemini-3-Flash-Preview with LightMem, AdaptiveMem recovers 14.9 percentage points on MemTrapBench while maintaining performance on LongMemEval.

Benchmark

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

Ryuichi Sumida, Koji Inoue, Tatsuya Kawahara

arXiv 2026 · 2026

MemUse combines a 4-month GPT-4.1-mini deployment with components like Summary-only, LC-k%, RAG-k%, and the MEMUSE benchmark to study memory use in real conversations. MemUse’s main result shows Direct QA rising from 19.7% to 70.1% across conditions while satisfaction stays flat, and within MEMUSE moments Natural Integration, not Direct QA, is associated with higher user satisfaction.

Benchmark

WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems

Jiangnan Yu, Kisson Songqi Lin, Jilong Wu

arXiv 2026 · 2026

WhenLoss evaluates long-context memory systems by running a fixed reader under Truncated Full Context, Oracle Evidence, Complete Stored Memory, and Retrieved Memory to expose write and retrieval bottlenecks. On 500 LongMemEval questions, WhenLoss with Expected Predictive Compression reaches 0.49 CSM Contains Match vs 0.44 for Summary LLM, reducing the write-side gap to 0.04.

BenchmarkLong-Term Memory

Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs

Mohammad Tavakoli, Alireza Salemi et al.

arXiv 2025 · 2025

LIGHT augments LLMs with Retrieval from the Conversation, Scratchpad Formation and Utilization, and a Working Memory buffer plus noise filtering to answer BEAM’s long-context probing questions. On the BEAM benchmark, LIGHT raises GPT-4.1-nano’s average score at 10M-token conversations from 0.109 to 0.226, a +107.3% gain over the vanilla long-context baseline.

BenchmarkBenchmarkAgent MemoryMemory Architecture

Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents

Saad Alqithami

· 2025

MaRS organizes agent memory into episodic, semantic, social, and task nodes with provenance, scored by a privacy-aware retention controller and governed by FIFO, LRU, Priority Decay, Reflection-Summary, Random-Drop, and Hybrid policies. On the FiFA benchmark, the Hybrid policy in MaRS achieves a composite score of ≈0.911 across 300 runs and five memory budgets, outperforming simpler policies while preserving privacy and cost efficiency.

BenchmarkBenchmark

MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems

Qingyao Ai, Yichen Tang et al.

arXiv 2025 · 2025

MemoryBench orchestrates Task Provider, User Simulator, and Performance Monitor to feed heterogeneous tasks, simulate explicit and implicit feedback, and score LLM systems across declarative and procedural memory. MemoryBench’s main finding is that state-of-the-art memory systems like A-Mem, Mem0, and MemoryOS often fail to beat naive BM25 or embedding-based RAG on partitions such as SiLo and LiLo.

Benchmark

LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

Di Wu, Hongwei Wang et al.

ICLR 2025 · 2024

LongMemEval evaluates long-term interactive memory by running chat assistants through indexing, retrieval, and reading over 50k sessions with fact-augmented keys and time-aware query expansion. On LONGMEMEVALS, long-context LLMs like GPT-4o, Llama 3.1, and Phi-3 suffer 30%–60% accuracy drops compared to oracle evidence-only reading, revealing severe limitations in current long-context designs.