Reasoning Models Lose their Way with Long Deduction
Hadeel Alnegheimish ⋅ Jasna Ilieva ⋅ Yoon Kim
Abstract
Current frontier LLMs can theoretically process long contexts with up to $1M$ tokens. But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts? We propose a simple synthetic benchmark to probe for long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog. Our benchmark, dubbed ProloNg, systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of $22$ and $62k$ context length. We study 8 reasoning models across 5 families of frontier LLMs, and find that all models fail to perform above chance as the reasoning depth grows, with the majority failing beyond depth-$10$.
Chat is not available.
Successful Page Load