Calibrate Once, Choose the Beam: A Predictive Regime Test for Same-LM Search Guidance and Pruning
Jiahui Qu ⋅ Yifang Qin ⋅ BOYANG ZHENG ⋅ Ziyi Zhou
Abstract
We show that one calibration scalar decides whether same-LM beam search beats independent sampling and at what beam width, before any deployment-time sweep is run. That scalar is $p$, the precision with which the deployed LM scores its own partial plans as a search-pruning verifier, measured once on a labeled split (24 mazes, 160 labeled successors). A frontier-survival argument turns that one number into a regime test: a pre-deployment check on whether to run search and at what width. The test takes the form $A_k \approx [1-(1-p)^k]^{\bar L}$, whose slack points in the conservative direction under positive step-to-step error correlation. Plugged blindly into a disjoint 50-maze holdout, the equation correctly places the useful-$k$ regime for both Qwen2.5-7B ($p{=}0.816$) and Qwen2.5-3B ($p{=}0.755$), giving $k^\star{\in}\{3,4\}$ at $\tau{=}0.95$. On the same A100, self-consistency (SC-10) solves 9\% while same-LM guided $k{=}3$ solves 98\% at lower wall-clock; the ordering guided $>$ sampling additionally holds, as a qualitative transfer probe, under a visual-language modality swap and 100-game Gomoku self-play. For LM4Plan practitioners, this reduces two design decisions (whether to run search at all, and how wide) to one calibration measurement, rendered as a design map from $(p,\bar L)$ to the minimum admissible beam $k^\star$.
Chat is not available.
Successful Page Load