TerminalBench
TerminalBench is a benchmark tool designed to evaluate AI agents operating within a terminal environment, focusing specifically on agent behavior and performance in command-line contexts.
Overview
Rustem Feyzkhanov of Snorkel AI identifies TerminalBench as a benchmark focused on "agent running in terminal" — distinguishing it from broader or more abstract evals by grounding assessment in the concrete, tool-using context of a terminal interface. ↗
No further detail about TerminalBench's task structure, scoring methodology, or relationship to other benchmarks is available from the current material.