MADE: Benchmark Environments for Closed-Loop Materials Discovery
Abstract
Existing benchmarks for computational materials discovery primarily evaluate static predictive tasks or isolated computational sub-tasks. While valuable, these evaluations neglect the inherently iterative and adaptive nature of scientific discovery. We introduce MAterials Discovery Environments (MADE), a novel framework for benchmarking end-to-end autonomous materials discovery pipelines. MADE simulates closed loop discovery campaigns in which an agent or algorithm proposes, evaluates, and refines candidate materials under a constrained oracle budget, capturing the sequential and resource-limited nature of real discovery workflows. We formalize discovery as a search for thermodynamically stable compounds relative to a given convex hull, and evaluate efficacy and efficiency via comparison to baseline algorithms. The framework is flexible; users can compose discovery agents from interchangeable components such as generative models, filters, and planners, enabling the study of arbitrary workflows ranging from fixed pipelines to agentic systems. We demonstrate this by conducting systematic experiments across a diverse range of systems and algorithms, finding that adaptive planning becomes more important to discovery efficiency as the search space scales.
Lay Summary
Scientific discovery is a closed-loop process. Researchers propose hypotheses, run experiments or simulations, and refine their ideas based on the outcomes. However, benchmarks evaluating algorithms at automated materials discovery do not capture this iterative, feedback-driven process. We propose a flexible and modular framework for evaluating AI scientists and algorithms at end-to-end closed-loop computational materials discovery. This enables practitioners to simulate discovery campaigns and effectively measure and compare the performance of algorithms before using them on expensive real-world experiments. We show that pipelines that adapt strategies given feedback are significantly more effective at discovery than static pipelines, particularly when the search space becomes more complex. This highlights the need for benchmark environments that evaluate closed-loop discovery.