Evaluation map

AI memory benchmarks

Compare the benchmarks used to evaluate conversational, long-context, and agent memory systems, with links to their original papers and datasets.

7 verified benchmarksOriginal papers and artifactsLimitations included

Compare the field

Choose a benchmark by what you need to test

BenchmarkBest forEvaluationMemory horizonReleased
DolphinBenchLong-horizon memory-guided tool useTool-using agent tasks in simulated appsNearly five years of conversation history2026
LoCoMoLong-term conversational recall and reasoningConversation QA and generationUp to 35 sessions2024
BEAMExtreme-length conversational memoryLong-context conversation QAUp to 10M tokens2025
MemoryArenaMemory-guided agent actionsInteractive agent tasksInterdependent multi-session tasks2026
LongMemEvalLong-term assistant memoryMulti-session conversation QAScalable histories; 500 questions2024
MemBenchMulti-dimensional agent-memory evaluationAgent interaction scenariosFactual and reflective memory2025
MemoryBenchLearning continuously from user feedbackSimulated service-time feedbackStepwise continual learning2025

Benchmark atlas

Understand each benchmark before comparing scores

Benchmark

DolphinBench

An open benchmark for long-term agent memory that evaluates whether a system uses years of conversation history to take the correct tool action, while also reporting cost and latency.

600 verified tasks across three simulated users with roughly 500K tokens of message history per persona.

Explore DolphinBench →Official project ↗

Benchmark

LoCoMo

A dataset and evaluation benchmark for very long-term conversational memory across question answering, event summarization, and multimodal dialogue generation.

Very long conversations averaging 300 turns and 9K tokens across as many as 35 sessions.

Explore LoCoMo →Official project ↗

Benchmark

BEAM

A long-context conversational-memory benchmark with coherent conversations reaching 10M tokens and questions spanning multiple memory abilities.

100 conversations, up to 10M tokens each, with 2,000 human-validated questions.

Explore BEAM →Official project ↗

Benchmark

MemoryArena

An evaluation gym for agent memory in interdependent multi-session tasks where agents must learn from earlier actions and reuse that memory later.

Human-authored, interdependent task sequences evaluated through Memory-Agent-Environment loops.

Explore MemoryArena →

Benchmark

LongMemEval

A scalable benchmark for long-term interactive memory in chat assistants, covering information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention.

500 curated questions across five long-term memory abilities, with histories that can be scaled in length.

Explore LongMemEval →Official project ↗

Benchmark

MemBench

A benchmark for LLM-agent memory across factual and reflective memory, participation and observation scenarios, and effectiveness, efficiency, and capacity dimensions.

Multiple memory levels, interaction scenarios, and evaluation dimensions in one agent-memory suite.

Explore MemBench →

Benchmark

MemoryBench

A benchmark for memory and continual learning in LLM systems that tests declarative and procedural memory learned from explicit and implicit user feedback.

28 benchmarks across three domains, four task formats, and two languages, with eight published baselines.

Explore MemoryBench →Official project ↗

How to use this atlas

A benchmark score is not a universal memory score

Match the workload

Use conversational QA benchmarks for recall, interactive suites for tool-using agents, and feedback benchmarks for continual learning.

Check the protocol

Model, judge, retrieval depth, prompt, data version, and cost accounting can make apparently similar scores incomparable.

Read limitations

Every atlas page names what the benchmark does not test, so a leaderboard result is not mistaken for production readiness.