BARSA: An Adaptive Test-Time Scaling Strategy for Mathematical Reasoning under Global Compute Budgets
Yufan Zhao ⋅ Yinsicheng Jiang ⋅ Cheng Deng ⋅ Yeqi Huang ⋅ Tairan Xu ⋅ Zhan Lu ⋅ Luo Mai ⋅ Wenda Li
Abstract
This paper presents our submission to the AI Mathematical Olympiad - Progress Prize 3 (AIMO 3) competition. We propose **Budget-Aware Recursive Self-Aggregation (BARSA)**, an adaptive test-time scaling framework for mathematical reasoning under a global inference budget. BARSA extends Recursive Self-Aggregation by using answer-distribution statistics and runtime estimates to decide whether to accept the current answer or continue with another aggregation round. We evaluate BARSA on the AIMO 3 leaderboards and a curated benchmark of AI-hard problems. Our analysis shows that recursive aggregation is most effective when the correct answer appears as a minority candidate or when the answer distribution is unstable, but remains vulnerable to deceptive wrong majorities. BARSA improves the mean accuracy and reduces score variance when evaluated under a fixed time budget. Across public leaderboard submissions, BARSA achieved a mean score of $40.38$ with standard deviation $0.744$, compared with $36.70 \pm 1.567$ for Majority@8 with a dynamic scheduler and $37.91 \pm 1.676$ for RSA with the same dynamic scheduler. These observational leaderboard results suggest that coupling recursive aggregation with budget-aware scheduling can improve average performance and may improve run-to-run stability.
Chat is not available.
Successful Page Load