Bayesian Frontier-Evaluation Testing Under Repeated Top-k Reporting
Yanan Long
Abstract
Frontier evaluation is a sequential decision problem under selective reporting. Public evidence often arrives as repeated top-$k$ snapshots rather than full score distributions, so the primitive object is not rejection of a null hypothesis but a frontier-evaluation action chosen under asymmetric loss. We formulate Bayesian testing for this setting by combining a selection-aware frontier model with hypotheses defined as regions of a latent state space. The mathematical core is an observability result: repeated snapshots can constrain a common frontier trajectory, whereas terminal-only archives can agree on the selected terminal law while leaving action-relevant timing unidentified. Synthetic studies illustrate the decision layer while supporting only bounded claims. A repeated-snapshot protocol study covers $225$ fitted records and obtains $219/225$ correct binary horizon decisions, while a posterior-decision analysis covers $374$ fitted records and shows substantially lower frontier-path loss for the repeated-snapshot model. These results are deliberately limited: plateau-time interval coverage remains weak, and the matched terminal-only comparator ties the repeated-snapshot model on the binary horizon decision in the full matched comparison.
Chat is not available.
Successful Page Load