DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven Evolution
Abstract
LLM–driven evolutionary systems have shown promise for automated science discovery, yet existing approaches such as AlphaEvolve rely on full-code histories that are context-inefficient and potentially provide weak evolutionary guidance. In this work, we first formalize the evolutionary agents as a general Expectation–Maximization framework, where the language model samples candidate programs (E-step) and the system updates the control context based on evaluation feedback (M-step). Under this view, constructing context via full-code snapshots constitutes a suboptimal M-step, as redundant implement details dilutes core algorithmic ideas, making it difficult to provide clear inspirations for evolution. To address this, we propose DeltaEvolve, a momentum-driven evolutionary framework that replaces full-code history with structured semantic delta capturing how and why modifications between successive nodes affect performance. As programs are often decomposable, semantic delta usually contains many effective components which are transferable and more informative to drive improvement. By organizing semantic delta through multi-level database and progressive disclosure mechanism, input tokens are further reduced. Empirical evaluations on tasks across diverse scientific domains show that our framework can discover better solution with less token consumption over full-code-based evolutionary agents.
Lay Summary
Large language models can now help discover new algorithms by writing many candidate programs, testing them, and iteratively improving on the best ones. However, existing systems are highly inefficient: every time the model proposes a new program, it must reread the full source code of past programs for inspiration, which wastes computation and obscures which past changes actually led to improvements. We introduce DeltaEvolve, which replaces these bulky code snapshots with short notes describing only what changed between successive programs and why that change helped or hurt. This gives the model a longer, cleaner memory of useful ideas while reading far less text. Across five scientific problems — including geometric packing, equation discovery, and equation solving — DeltaEvolve finds better solutions while using roughly 37% fewer input tokens than the leading prior method, making automated scientific discovery substantially cheaper and easier to scale.