Implications of Large-Scale Test-Time Compute
Noam Brown
Abstract
As LLMs become more capable, their performance increasingly depends on the amount of test-time compute used. This talk argues that single-number benchmarks obscure both capability and safety-relevant trends, especially as longer chains of thought, scaffolds, and multi-agent methods push performance higher with more inference. I will close with proposals for evaluating models using performance-vs-compute curves and updating benchmarks and preparedness frameworks to account for realistic high-compute use.
Speaker
Noam Brown
Video
Chat is not available.
Successful Page Load