Shifting a Molecular Generator Toward Developability with Iterative Importance Fine-Tuning
Abstract
Designing small molecules with strong target binding while maintaining favorable absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties or synthetic accessibility remains a central challenge in drug discovery. Recent generative models such as GenMol enable efficient exploration of chemical space for hit identification and lead optimization. However, optimizing for a single objective (e.g., protein binding) often degrades other critical ADMET or synthesizability properties. We propose a framework that combines GenMol with Iterative Importance Fine-Tuning (IIFT), a reward-based post-training method that shifts the generative distribution toward molecules satisfying multiple objectives. IIFT requires only a scalar reward and does not assume differentiability, allowing incorporation of black-box oracle functions. We employ IIFT to develop two distribution-shifted variants of GenMol: \textbf{Viable GenMol}, which biases generation toward ADMET-favorable molecules, and \textbf{Tractable GenMol}, which biases generation toward synthetically accessible molecules. We evaluate these models in \textit{de novo} generation, hit identification, and lead optimization, achieving replicable improvements across all three modalities while preserving drug-likeness, synthetic accessibility, and molecular diversity. We validate selected tractable lead-optimized candidates using binding free energy calculations of molecular affinity. Our results demonstrate that importance-based distribution shifting provides a practical approach to multi-objective molecular design under realistic drug discovery constraints.