Dynamic large language model representations for multi-objective chemical reaction optimisation
Abstract
Optimising chemical reactions over multiple objectives, such as yield, selectivity, and cost, remains a central challenge in chemical synthesis. Current model-driven approaches rely on fixed molecular representations such as one-hot encodings that discard chemical knowledge, or molecular descriptors that require domain-specific selection workflows and lack common representations across chemically distinct reaction components. Here, we present a multi-objective Bayesian optimisation framework that replaces fixed featurisations with trainable large language model (LLM) embeddings. Textual descriptions of reaction conditions are encoded by a fine-tuned language model jointly optimised with Gaussian process surrogates, producing dense, task-adaptive representations that improve as more experimental data is collected. We benchmark our approach against established descriptor libraries and one-hot encoding baselines across sequential low-data optimisation of nickel- and palladium-catalysed cross-couplings and large-batch 96-well plate high-throughput experimentation regimes, achieving practical convergence thresholds in fewer iterations across all settings. Our framework provides a general, out-of-the-box approach to multi-objective reaction optimisation broadly applicable across reaction types and reaction components, without requiring descriptor selection workflows or domain-specific feature engineering.