SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
Abstract
Lay Summary
Can AI predict what will happen in a science experiment before it's actually run? This paper introduces SciPredict, a test of 405 questions drawn from real, recently published experiments in physics, biology, and chemistry to find out. The results are humbling: the best AI models only get about 14–26% of predictions right, roughly on par with human experts (\~20%). But the bigger problem isn't just accuracy — it's that AI models can't tell when they're guessing well versus guessing badly. They report similar confidence whether they're right or wrong. Human experts, by contrast, have good instincts about which questions they can answer without running the experiment (reaching \~80% accuracy on those) and which ones they can't (\~5%). The study also found that giving AI relevant background information helps a little, but when models try to come up with their own background knowledge, it actually makes their predictions worse. The takeaway: AI isn't yet reliable enough to guide scientists on which experiments are worth pursuing, and the missing piece isn't just smarter predictions — it's knowing when to trust those predictions.