A recipe for scalable attention-based ML potentials: unlocking long-range accuracy with all-to-all node attention
Abstract
Machine-learning interatomic potentials (MLIPs) have advanced rapidly, with many top models relying on strong physics-based inductive bias. However, as models scale to larger systems like biomolecules and electrolytes, they struggle to accurately capture long-range (LR) interactions, leading current approaches to rely on explicit physics-based terms or components. In this work, we propose AllScAIP, a straightforward, attention-based, and energy-conserving MLIP model that scales to O(100 million) training samples. It addresses the long-range challenge using an all-to-all node attention component that is purely data-driven. Extensive ablations reveal that in low-data/small-model regimes, inductive biases improve sample efficiency. However, as data and model size scale, these benefits diminish or even reverse, while all-to-all attention remains critical for capturing LR interactions. Our model achieves state-of-the-art energy/force accuracy on molecular systems (OMol25), while being competitive on materials (OMat24) and catalysts (OC20). Furthermore, it enables stable, long-timescale MD simulations that accurately recover experimental observables, including density and heat of vaporization predictions.
Lay Summary
Computer simulations can help scientists understand molecules and materials without running as many expensive laboratory experiments. A major challenge is building AI models that can accurately predict how atoms interact, especially when important effects come from atoms that are far apart. Many current models handle this by adding hand-designed physics rules. In this work, we introduce AllScAIP, a simpler AI model that learns these long-distance atomic interactions directly from data. The key idea is to let every atom in a system share information with every other atom, while still keeping efficient local intreactions for nearby atoms. AllScAIP achieves strong accuracy on large molecule datasets, performs competitively on materials and catalyst datasets, and can run stable simulations that reproduce experimental properties such as liquid density and heat of vaporization.