Benchmark
LoCoMo
A dataset and evaluation benchmark for very long-term conversational memory across question answering, event summarization, and multimodal dialogue generation.
Very long conversations averaging 300 turns and 9K tokens across as many as 35 sessions.
Explore LoCoMo →Benchmark
BEAM
A long-context conversational-memory benchmark with coherent conversations reaching 10M tokens and questions spanning multiple memory abilities.
100 conversations, up to 10M tokens each, with 2,000 human-validated questions.
Explore BEAM →Benchmark
MemoryArena
An evaluation gym for agent memory in interdependent multi-session tasks where agents must learn from earlier actions and reuse that memory later.
Human-authored, interdependent task sequences evaluated through Memory-Agent-Environment loops.
Explore MemoryArena →