Learning to Explore with Parameter-Space Noise: A Deep Dive into Parameter-Space Noise for Reinforcement Learning with Verifiable Rewards
Bizhe Bai ⋅ Xinyue Wang ⋅ Peng Ye ⋅ Tao Chen
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) enhances large language model (LLM) reasoning, yet growing evidence indicates an \emph{exploration ceiling}: they tend to reweight existing solution traces and produce low-diversity rollouts. This restricts the discovery of novel solution perspective and weakens out-of-distribution generalization. We address this limitation by directly perturbing the rollout policy's parameters while updating the clean policy model as usual, an approach that achieves better temporal consistency and yields more complex behavior patterns. To stabilize learning from these perturbed rollouts, which are inherently off-policy, we incorporate truncated importance sampling. Furthermore, we introduce a lightweight, dynamic noise scheduler based on semantic diversity and model self-certainty to optimally adjust the noise scale. Instantiated on GRPO, our method (PSN-GRPO) demonstrates its efficacy through four key contributions: (1) it improves high-budget pass@$k$, increases semantic and operational diversity, increases CoT consistency, and remains complementary to exploration-oriented RLVR techniques; (2) it validates efficacy across multiple models, including Qwen2.5, Qwen3, and Llama-3.1; (3) it shows robust performance on out-of-distribution science and coding benchmarks; and (4) it uncovers novel solution perspectives for problems the original model could not solve, as shown by qualitative analysis.
Chat is not available.
Successful Page Load