UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon Scenarios
Abstract
Lay Summary
While artificial intelligence has made remarkable progress on short, straightforward tasks, it still struggles with the messy, long-term challenges we face in the real world, such as scientific research or complex business investing. To test these limits, we created UltraHorizon, a massive new benchmark where AI agents must act like scientists to uncover hidden rules through continuous trial and error, often processing hundreds of thousands of words and making hundreds of decisions in a single attempt. Our experiments reveal that even the most advanced AI models fall far behind human participants in these scenarios. Furthermore, simply giving the AI more time or a larger interaction budget does not fix the problem. By analyzing their behavior, we found that AI agents typically fail because they get stubbornly stuck on a flaw we call "in-context locking", and because they lack the basic ability to reliably manage their memory and reasoning over long periods. Ultimately, UltraHorizon highlights a critical gap in current AI capabilities and provides a roadmap for building smarter, more persistent AI systems.