PRISM: Demystifying Retention and Interaction in Mid-Training
Abstract
Lay Summary
Modern AI language models are trained in stages. Researchers have recently added a "mid-training" step , a focused practice phase between initial learning and final fine-tuning but nobody had studied it carefully enough to know what actually works. We investigated this gap systematically across multiple AI models. We found that feeding models a carefully balanced diet of math, coding, and science problems during mid-training dramatically sharpens their reasoning: improving math scores by up to 30 points and coding scores by up to 10 points while preserving everything they already knew. Crucially, we also showed that this mid-training phase makes subsequent reinforcement learning far more effective and stable. Think of it as ensuring athletes master fundamentals before advanced coaching skipping this step leaves significant performance on the table. Our findings provide practical guidance for building more capable and better reasoning models.