Reasoning Quality Emerges Early: Data Curation for Reasoning Models
Abstract
Supervised fine-tuning (SFT) on a small, high-quality set of long reasoning traces is an effective approach for eliciting strong reasoning capabilities in Large Language Models (LLMs). However, existing methods for curating high-quality SFT data rely heavily on strong reasoning models to filter examples based on diversity and difficulty, making the curation process costly while often yielding suboptimal data quality. In this work, we show that diverse and challenging reasoning examples can be identified using only the initial reasoning tokens. Specifically, we demonstrate that difficult problems can be reliably detected based on the loss of the first 100 reasoning tokens evaluated at a randomly perturbed checkpoint of the pretrained model. We further show that examples exhibiting similar loss patterns over their first 1k reasoning tokens across a small number of perturbed checkpoints extrapolating along the fine-tuning trajectory provably induce similar gradients. We validate our approach through extensive experiments on fine-tuning Qwen2.5-7B and Llama3.1-8B models on the M23K medical reasoning and OpenThoughts-Math datasets. Our method outperforms existing baselines by up to 1.7% while being 91% more token efficient.
Lay Summary
AI models can think and reason when solving complex problems when they are trained on a hard and diverse set of problems, but selecting these data is prohibitively expensive because existing methods need to read through every single word in all the examples before picking the best ones. We show that we can add noise to the AI model's parameters and have it read the first 100 to 1,000 tokens of each example response, and this is sufficient to measure the problem's difficulty and diversity. This allows us to select the best training examples without reading the full text. Our Token-Efficient Model Perturbation (TEMP) technique cuts the amount of processed text by 91% and boosts the final model's reasoning accuracy by up to 1.7%.