AMR-Bench-mini: A Diagnostic Benchmark for Agentic Mechanistic AMR Reasoning under Evidence Insufficiency
Abstract
Artificial intelligence agents powered by frontier large language models (LLMs) have rapidly converged on a planner+tools+critic (PTC) architecture for biological reasoning, but it is still not clear whether they reason about evidence sufficiency or merely pattern-match plausible mechanisms. We introduce AMR-Bench-mini, a focused diagnostic benchmark for agentic mechanistic antimicrobial-resistance (AMR) reasoning, scoped strictly to retrospective evidence synthesis on Klebsiella pneumoniae. The benchmark pairs a 924-task corpus generated from public isolate metadata, genome assemblies, laboratory AST, and pinned annotation outputs with two evaluation splits: a balanced 50-task pilot and a 50-task hard diagnostic split organised into seven category-level failure modes. We document an audit-and-fix loop that lifts heuristic gold credibility from 71 % to 93 %, with four disagreements where all three vendors converged against the heuristic gold. Three frontier models (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.7) score 82–88 % on the pilot and 50–64 % on the hard split. All three converge on∼25 % accuracy on the abstention category, while contrastive-control categories remain high descriptively. A pre-resolved CARD ARO substrate-context tool injected into Gemini’s prompt produces a zero-percentage-point net effect on hard accuracy, with two gains and two losses: substrate retrieval is necessary but not sufficient for evidence-sufficiency reasoning.