Best for
- • Long-horizon memory-guided tool use
Memory benchmark
An open benchmark for long-term agent memory that evaluates whether a system uses years of conversation history to take the correct tool action, while also reporting cost and latency.
Original research
This is the paper that introduced DolphinBench.
arXiv:2609.24971 ↗