Benchmarks as Software: A Case Study on Terminal-Bench
Estefany Kelly Buchanan ⋅ Scott Linderman ⋅ Christopher Re
Abstract
Agentic benchmarks accumulate defects on every axis: instructions can be overspecified, verifiers can be underspecified, and environments drift as external dependencies change after release. We audit Terminal-Bench~2.0 with a continuous-validation pipeline (automated CI checks, LLM-as-judge). We find task-side defects in 28 of 89 tasks (31\%) in the benchmark. We ship the fixes as Terminal-Bench~2.1. Across 19 agents, scores rise by up to 12~pp, but agent ranking is stable (Spearman $\rho = 0.958$). On the changed tasks, 28.8\% of the tokens TB~2.0 charged to failed trials on changed tasks were spent on what TB~2.1 verifies as solvable work. Cost-aware curves show that the audit shifts per-agent efficiency in both directions: some agents gain coverage while becoming more token-efficient, others while becoming less. On TB~2.1 benchmark, switching from \texttt{codex} to \texttt{terminus-2} at fixed OpenAI model drops accuracy by up to 29.9 pp; raising the per-task budget closes the gap, indicating the harness effect is wall-clock-bound rather than capability-bound.
Chat is not available.
Successful Page Load