Synthetic Data and the Rise of Spiky Intelligence
Abitha Thankaraj ⋅ Amro Abbas ⋅ Dongyang Fan ⋅ Vineeth Dorna ⋅ Luke Merrick ⋅ David Schwab ⋅ Anshuman Suri ⋅ Aldo Carranza ⋅ Alex Fang ⋅ Alvin Deng ⋅ Brett Larsen ⋅ Darren Teh ⋅ Diego Kiner ⋅ Fan Pan ⋅ Haakon Mongstad ⋅ Haoli Yin ⋅ Jack Urbanek ⋅ Jason C Lee ⋅ Jason Telanoff ⋅ Josh Wills ⋅ Katherine L Mentzer ⋅ Maximilian Böther ⋅ Parth Doshi ⋅ Paul Burstein ⋅ Rishabh Adiga ⋅ Siddharth Joshi ⋅ Tony Jiang ⋅ Vidhi Jain ⋅ Zhengping Wang ⋅ Yonatan Bisk ⋅ Bogdan Gaza ⋅ Ari Morcos ⋅ Matthew Leavitt ⋅ Pratyush Maini
Abstract
Synthetic data is increasingly used to train language models, but standard decontamination methods mainly detect duplicated or near-duplicated test examples. They miss a subtler failure mode: \textit{benchmaxxing}, where training data is synthetic and benchmark-adjacent without containing the test set. We study benchmaxxing under controlled conditions and compare it with direct test-set contamination across model scales. We introduce the \texttt{Inflation Score}, which measures how much augmentation increases the benchmark--held-out accuracy gap relative to a no-augmentation baseline, using a capability-matched held-out set. We find that benchmark-adjacent synthetic data can inflate benchmark scores as much as direct test-set training. To test transfer, we introduce \texttt{Factory-GSM}, a GSM8K rewrite that preserves solution structure while changing surface context. Benchmaxxed models degrade substantially on \texttt{Factory-GSM}, suggesting benchmark-specific shortcuts rather than robust reasoning. Gradient similarity detects this specialization better than embedding-based decontamination, though imperfectly. A real-world audit of $46$ open-weight models shows similar patterns. Overall, reliable evaluation must detect concentrated benchmark-adjacent training signals, not just copied test examples.
Chat is not available.
Successful Page Load