Calibrate Once, Choose the Beam: A Pre-Deployment Compute-Allocation Rule for Same-LM Multimodal Search
Jiahui Qu ⋅ Yifang Qin ⋅ BOYANG ZHENG ⋅ Ziyi Zhou
Abstract
Agentic planners must decide whether a fixed inference budget should buy more complete samples or a managed search frontier. We give a pre-deployment rule for this choice. The deployed LM, used as its own pruning verifier, is calibrated once to a precision $p$ on locally true-improving partial states. Combining $p$ with the mean horizon $\bar{L}$ yields a frontier-survival score $A_k \approx [1 - (1-p)^k]^{\bar{L}}$, and the smallest beam $k$ whose score clears a target success rate is the recommended width. The oracle labels used to measure $p$ appear only on a small calibration split, never inside the test-time solver. On our 100-maze MazeBench (4×4 to 6×6, generator-controlled), the rule predicts the useful regime at $k=3$ from a single calibration scalar $p=0.816$. Same-LM guided search then solves 98 of 100 mazes in 14.4 seconds per maze on one A100, while SC-10 solves 9 in 17.7 seconds. A Qwen2.5-3B holdout, a Qwen2.5-VL-7B run on rendered visual mazes (40/50 vs 1/50 for SC-10), and a same-LM two-ply Gomoku tournament (100/100 wins) all land on the safe side of the calibrated boundary, yielding a design map from $(p, \bar{L})$ to the minimum admissible beam across model size, modality, and task.
Chat is not available.
Successful Page Load