Optimistic Online Learning for Data Mixture Optimization
Sofija Orlovic ⋅ Francesco Tonin ⋅ Volkan Cevher
Abstract
Data mixture optimization is important for language model pretraining, where models are trained on heterogeneous corpora with domain-dependent usefulness. We propose Optimistic DoReMi, a minimal modification of DoReMi that replaces its multiplicative domain-weight update with an optimistic update inspired by Optimistic Hedge. Instead of using only the current per-domain excess-loss vector, our method updates weights using a linear prediction from the two most recent estimates. Experiments on SlimPajama and the Pile datasets show that this simple modification yields more stable and accurate mixtures over DoReMi across scales.
Chat is not available.
Successful Page Load