Same Biology, Different Scores: Quantifying the Tool-Use Confound in Agentic Biology Evaluation
Abstract
Agentic biology benchmarks are increasingly used to assess AI capabilities in the life sciences. When a model fails an agentic bio task, current benchmarks cannot distinguish whether the failure reflects missing domain knowledge or an inability to execute code—a distinction that determines whether better tools or better training is the appropriate intervention. We introduce DISSECT (Decomposing In-Silico Scientific Evaluation into Capability and Tool-use), a 9-task agentic biosecurity benchmark with matched ablation pairs that test identical scientific content with and without code requirements—to our knowledge, the first such decomposition in agentic biology evaluation. Across two ablation pairs spanning molecular biology and biosecurity reasoning, evaluated on multiple frontier models, we find preliminary evidence that the confound may be substantial: models that score near-ceiling on no-code versions can collapse to near-floor when code execution is required, and capability rankings can flip depending on whether code is required, suggesting that "biology capability" as currently measured may not be a single dimension. We also observe that standard MCQ benchmarks do not distinguish safety-filtered refusals from genuine capability deficits—and that refusal behavior is context-dependent, with the presence of a tool interface triggering refusals that disappear when tools are removed. On reasoning-heavy tasks, code tools can be counterproductive. These findings, if they generalize, suggest that decomposing knowledge from tool-use is an important step toward the per-dimension capability assessments that evaluation, safety, and policy of biological AI agents increasingly require—and that current agentic bio benchmarks do not provide this decomposition by default.