Continual Learning Bench
A benchmark designed to evaluate AI agent performance on long-horizon tasks, specifically measuring how well agents maintain and apply knowledge across extended task sequences.
Overview
Continual Learning Bench is a benchmark created and run by Lance Martin at Anthropic to assess Claude agents on long-horizon tasks. Martin introduced it in the context of evaluating agent capabilities over sustained, multi-step workflows — a core challenge in building reliable AI systems that must retain context and performance across complex task horizons. Lance Martin, Claude for Long-Horizon Tasks, 3:21:10
Beyond this single mention, no further detail about the benchmark's methodology, metrics, or results is available in the current graph.