Conformal C2ST: Turning weak classifiers into strong two-sample tests
Abstract
Lay Summary
Scientists increasingly use AI models that imitate complex real-world processes — for instance, generating videos, text, or even reconstructing a galaxy's true appearance from light that was distorted on its way to our telescopes. But how do we know such an answer is genuinely right, rather than subtly wrong in a way that could mislead a scientific conclusion? A common check trains a second program to separate the model's outputs from real examples: if it cannot tell them apart, the model passes. The trouble is that this only works when that program is very accurate, and in practice it is often weak. A weak program can wrongly clear a flawed model, so a "pass" may simply mean the check was too poor to notice the problem — not that the model is sound. We show that even a weak program can be turned into a reliable, sensitive test. Instead of relying on its yes-or-no judgments, we use only the rough ordering it assigns to examples — something even a poor program gets partly right — and a simple post-processing converts these rankings into trustworthy evidence with guaranteed error rates. Our test catches subtle mistakes that other methods miss and stays dependable even when the program behind it is unreliable, giving scientists a practical safety check for the AI models they rely on.