OpenSeesAgentBench: A Benchmark and Evaluation Framework for Agentic Structural Analysis in OpenSeesPy
Abstract
Large language model (LLM) agents are being applied to structural analysis, but evaluation is often reduced to executability and can blur benchmark construction with agent capability. We present OpenSeesAgentBench, an OpenSeesPy benchmark and executable reference evaluation stack spanning OpenSeesQuery (90 questions), OpenSeesCodeBench (24 code tasks), and OpenSeesWorkflowBench via the OpenSeesBuildingBench calibration profile (900 workflow cases). The workflow layer uses contract-first evaluation with strict CaseSpec validation, fail-closed checks, and bounded repair. In the current paper, we report only the reference stack, so the workflow results should be read as a calibration study rather than a head-to-head systems comparison: observed shard expected-match rates of 89.33\%, 70.33\%, and 60.67\% stay within one case of the profile targets. These results show that the released workflow profile preserves its intended difficulty under executable expansion and provides an auditable benchmark for future comparative evaluation; they do not, by themselves, establish superiority of a particular agent architecture.