AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
Yiheng Shu, Bernal Jiménez Gutiérrez et al.
arXiv 2026 · 2026
AGENTCL evaluates continual learning in language agents by contrasting naive and compositional task streams with a two-pass protocol and MEMPROBE’s interaction memory, insight memory, and skill memory. On the CodeEval-Pro compositional stream, AGENTCL shows MEMPROBE achieving 66.7% first-pass accuracy versus the memoryless ReAct’s 48.8% (+17.9 pp), while naive streams show much smaller separations.