ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
Abstract
Lay Summary
Testing new AI takes a lot of time and money because developers usually have to check thousands of answers. To save cost, they often only test a small batch of questions, which means they can easily miss rare but dangerous mistakes. To fix this, we created ProEval, a smart system that uses past test results to guess where a new AI is most likely to mess up. Instead of testing randomly, ProEval handpicks the trickiest questions to intentionally expose the AI's weak spots. It can even ask another AI to invent brand-new, super-hard questions on the exact topics the new AI is already struggling with. This makes testing incredibly fast, requiring up to 65 times fewer questions to figure out how good an AI actually is. Plus, it uncovers up to 5 times more hidden flaws and safety issues than normal testing methods. Ultimately, this helps build safer, more trustworthy AI without breaking the bank.