Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
Abstract
Model merging has emerged as a lightweight paradigm for enhancing Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. In this work, we analyze late-stage pre-training trajectories and uncover a \textbf{Rank-1 Subspace} phenomenon: while raw optimization steps oscillate violently, consecutive \emph{merged} checkpoints collapse onto a stable, approximately one-dimensional linear manifold. We theoretically ground this observation in a \emph{river-valley} landscape analysis: averaging acts as a geometric low-pass filter that dampens high-curvature noise to reveal the optimal descent direction. Capitalizing on this insight, we propose \textbf{Extra-Merge}, a training-free strategy that extrapolates along this subspace to minimize loss without additional gradient updates. Extensive experiments across GPT-2 and LLaMA families (124M to 2B) demonstrate that Extra-Merge consistently outperforms standard merging baselines. Notably, it yields consistent zero-shot accuracy gains on Pythia-12B downstream tasks and generalizes effectively to the Muon optimizer (Jordan et al., 2024).
Lay Summary
This paper studies why combining different saved versions of a language model can improve it. It finds that although training changes can look messy, averaged model versions follow a simple and stable direction. Based on this, the paper proposes Extra-Merge, a method that moves further along this direction without extra training. Experiments on several model families show that it improves over standard merging methods, including on downstream tasks and with the Muon optimizer.