Benchmark
DolphinBench
An open benchmark for long-term agent memory that evaluates whether a system uses years of conversation history to take the correct tool action, while also reporting cost and latency.
600 verified tasks across three simulated users with roughly 500K tokens of message history per persona.
Explore DolphinBench →Official project ↗