Memory benchmark

DolphinBench

An open benchmark for long-term agent memory that evaluates whether a system uses years of conversation history to take the correct tool action, while also reporting cost and latency.

Best for

  • Long-horizon memory-guided tool use

Evaluation style

  • Tool-using agent tasks in simulated apps

Memory horizon

  • Nearly five years of conversation history

Scale

  • 600 verified tasks across three simulated users with roughly 500K tokens of message history per persona.

Tasks

  • Memory-guided tool use
  • Long-horizon rule application
  • Stateful application actions

What it measures

  • Task accuracy
  • Total cost
  • Median latency

Original research

DolphinBench: Mapping the Pareto Frontier of Agent Memory

This is the paper that introduced DolphinBench.

arXiv:2609.24971