Learning Reaction-Condition Plausibility for Evaluation and Self-Curation Under Noisy and Non-Unique Supervision
Abstract
Many scientific benchmarks are built on reference labels that are noisy, incomplete, or non-unique. For reaction-condition prediction, exact match can therefore mark a workable protocol as wrong when it differs from the archived record. We present UniCon, a framework for reaction-condition evaluation and self-curation that learns the plausibility of reaction-condition pairs instead of trying to reproduce a single archived label. UniCon aligns 2D graph representations of reaction transformations with fingerprint-based embeddings of chemical conditions in a shared latent space. It scores a candidate protocol against the archived one using a Bradley-Terry preference, which we call UniConScore. The same score can curate training data by comparing archived records with predictions from an independent condition prediction model and filtering chemically anomalous entries. Across expert preference studies, zero-shot high-throughput experimentation benchmarks, and self-curation analyses, UniConScore gives a ranking signal that tracks chemical viability better than exact match under noisy supervision.