CADEngBench: Can AI Systems Co-Author Engineering Designs? A Hierarchical Benchmark for Physics-Verified Parametric CAD Generation
Harmanjot Singh ⋅ Abhra Kanti Dubey
Abstract
AI systems are increasingly proposed as co-authors for engineering design, yet current CAD evaluations largely reward executable geometry rather than reusable, physics-valid parametric models. We introduce CADEngBench, a hierarchical benchmark of 40 mechanically meaningful CadQuery tasks spanning 12 engineering domains and 17 physics tags. CADEngBench evaluates generated designs across four levels: $L0$ compilation and topological validity, $L1$ geometric fidelity, $L2$ parametric integrity via perturbation robustness and $Z3$ SMT constraint verification, and $L3$ closed-form physics verification. Across five representative models evaluated with a standardized prompting and scoring pipeline, Claude Sonnet 4.6 achieves the strongest end-to-end performance at $22/40$ passes ($55.0\%$). The benchmark's most diagnostic result is a specialist-generalist inversion: CADCoder attains $39/40$ at $L0$ but only $2/40$ at $L2$ and $1/40$ overall, demonstrating that runnable CAD code is not equivalent to engineering-grade CAD. Prompt underspecification is catastrophic, with $1/60$ overall passes at $\mathrm{level}_1$ versus $20/60$ at $\mathrm{level}_3$, and repeated sampling improves frontier hosted models but not CADCoder. These findings identify $L2$ parametric integrity as the principal failure boundary for current systems and position CADEngBench as a benchmark for whether AI systems can co-author reusable, physics-validated engineering designs.
Chat is not available.
Successful Page Load