Continual Learning Bench

concept · updated Jul 25, 2026

A benchmark designed to evaluate AI agent performance on long-horizon tasks, specifically measuring how well agents maintain and apply knowledge across extended task sequences.

Overview

Continual Learning Bench is a benchmark created and run by Lance Martin at Anthropic to assess Claude agents on long-horizon tasks. Martin introduced it in the context of evaluating agent capabilities over sustained, multi-step workflows — a core challenge in building reliable AI systems that must retain context and performance across complex task horizons. Lance Martin, Claude for Long-Horizon Tasks, 3:21:10

Beyond this single mention, no further detail about the benchmark's methodology, metrics, or results is available in the current graph.