The Inference Gap: Test-Time Compute Scaling and the Limits of Training-Centric AI Governance
Abstract
Contemporary AI governance frameworks, including the EU Artificial Intelligence Act and frontier developers' responsible scaling policies, assume that AI capability is determined primarily at training time. Test-time compute (TTC) scaling, wherein the same model weights produce qualitatively different outputs depending on inference budget, exposes gaps in this assumption that existing frameworks do not address. We identify three governance gaps that result from this mismatch. The first is a threshold decoupling gap, through which developers can pair sub-threshold training runs with inference-time amplification to achieve frontier-level capabilities outside the scope of enhanced regulatory scrutiny. The second is an evaluation staleness gap, through which pre-deployment safety assessments conducted at fixed inference budgets fail to characterise capabilities that emerge under extended reasoning. The third is a reasoning opacity gap, through which hidden chain-of-thought processes undermine the observability requirements that auditable governance depends on, with the additional complication that visible chain-of-thought is not reliably faithful even where accessible. We further analyse how agentic deployment compounds all three gaps and propose a set of technically grounded reforms, including inference-budget-aware capability evaluations and reasoning trace logging for regulated applications.