Evaluation map

AI memory benchmarks

Compare the benchmarks used to evaluate conversational, long-context, and agent memory systems, with links to their original papers and datasets.

Benchmark

LoCoMo

A dataset and evaluation benchmark for very long-term conversational memory across question answering, event summarization, and multimodal dialogue generation.

Very long conversations averaging 300 turns and 9K tokens across as many as 35 sessions.

Explore LoCoMo

Benchmark

BEAM

A long-context conversational-memory benchmark with coherent conversations reaching 10M tokens and questions spanning multiple memory abilities.

100 conversations, up to 10M tokens each, with 2,000 human-validated questions.

Explore BEAM

Benchmark

MemoryArena

An evaluation gym for agent memory in interdependent multi-session tasks where agents must learn from earlier actions and reuse that memory later.

Human-authored, interdependent task sequences evaluated through Memory-Agent-Environment loops.

Explore MemoryArena