Learning Compressed Shape-Aware Molecular Representations for Virtual Screening
Abstract
Lay Summary
Finding new medicines often starts with searching huge digital libraries of molecules — sometimes billions of them — to find ones that have a similar 3D shape to a known drug. Shape matters because molecules that "fit" the same biological target tend to look similar in three dimensions. The problem is that comparing 3D shapes is very slow: every molecule can fold into many possible 3D arrangements, and each one must be rotated and matched against the query. With today's libraries growing past ten billion compounds, this kind of search has become practically impossible. We trained a neural network called SAND that learns to predict 3D shape similarity directly from a molecule's flat 2D chemical structure, skipping the expensive 3D calculations entirely. We also taught it to compress each molecule into a small digital fingerprint that can be searched almost instantly. The result: searching ten billion molecules now takes less than a second on a single computer — over a hundred million times faster than traditional tools — making large-scale shape-based drug discovery more accessible.