Best for
- • Learning continuously from user feedback
Memory benchmark
A benchmark for memory and continual learning in LLM systems that tests declarative and procedural memory learned from explicit and implicit user feedback.
Original research
This is the paper that introduced MemoryBench.
arXiv:2510.17281 ↗Used in the field
Samuel Sameer Tanguturi
· 2026
ATANT v1.1 structurally analyzes seven benchmarks using the 7 v1.0 continuity properties, the 10 checkpoints, a property-coverage matrix, and the Kenotic v1.0 reference implementation. ATANT v1.1 reports 96% ATANT cumulative-scale versus 8.8% LOCOMO substring accuracy, showing that LOCOMO, LongMemEval, BEAM, MemoryBench, Zep eval, MemGPT/Letta, and RULER measure different properties from continuity.
Yihao Wang, Haoran Xu et al.
arXiv 2026 · 2026
MedMemoryBench builds a four stage pipeline with Patient Profile Construction, Disease Progression Event Generation, Multi turn Sessions Simulation, and Memory Extraction and Query Construction to synthesize long horizon, clinically grounded interactions. On MedMemoryBench Efficient vs Mixed, methods like Letta drop from 51.21% to 41.55% average accuracy, quantifying how memory saturation harms personalized healthcare agents.