AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
Abstract
Lay Summary
Problem: We often judge AI models by their scores on competitive leaderboards, but these rankings can be deceptive. A high score doesn’t always mean an AI is genuinely smarter; the results can be skewed by the model’s size, its specific training, or even quirks in the test itself. This makes it difficult to separate true progress from noise, leading researchers to chase misleading metrics. Solution: Our paper builds an “AI cartography” map that separates broad ability from benchmark-specific quirks and ecosystem effects (such as model size and common training choices). Using measurement tools designed to tease signal from noise, we estimate how much of each benchmark reflects general capability versus idiosyncrasies. This allows us to map the hidden landscape of AI capabilities, disentangling a model’s general intelligence from other influences so tiny leaderboard gaps aren’t over-interpreted. Impact: This new map provides a much clearer picture of AI progress. The map suggests knowledge-heavy tests keep improving strongly with scale, while gains on other reasoning-focused tests show much smaller returns once you control for general ability. Overall, the work helps the community make fairer, more meaningful comparisons—and reduces the risk of chasing misleading leaderboard wins instead of real advances.