CRYSTAL: Coordinated Multi-Objective Reinforcement Learning for Crystal Generation
Abstract
In materials science, \emph{de novo} generation of crystal structures that simultaneously satisfy stability, diversity, novelty, and validity is a central challenge for accelerating new materials discovery.Large language models, leveraging a two-stage paradigm of fine-tuning and reinforcement learning with verifiable rewards (RLVR), have demonstrated significant advantages over conventional diffusion-based approaches.However, existing RLVR methods typically constrain only one or two objectives, making it difficult to coordinate multi-objective optimization and prone to reward hacking---where one objective collapses sharply during training, leading to severe imbalance in overall performance.To address this, we propose CRYSTAL, a method based on Group Relative Policy Optimization that jointly optimizes four critical attributes: leveraging physics-grounded verifiable signals to ensure structural stability and physical correctness, regulating novelty through explicit comparison with established materials databases, and generating multiple candidate structures within a single inference step to explicitly promote diversity, ultimately achieving coordinated multi-objective control through multiplicative aggregation.The proposed method effectively mitigates reward hacking in multi-objective reinforcement learning, achieves state-of-the-art performance in comprehensive multi-objective evaluations, and attains an S.U.N.\ metric of 25.1 on 1,000 generated materials, demonstrating the potential to further extend large language models toward multi-objective on-demand materials design.