ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
Abstract
Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by model progress. This creates a need for difficult benchmarks that remain relevant for longer. We introduce ZeroBench—a lightweight visual reasoning benchmark curated using adversarial filtering to be “impossible” for frontier LMMs at its original release, with initial SotA scores of 0% pass@1 and pass∧5. We track progress on ZeroBench over the subsequent year, observing SotA reaching 6% pass∧5 and 19% pass@5, indicating the potential longevity of the benchmark. We evaluate 46 LMMs on ZeroBench, compare performance to a human baseline, analyse strengths and weaknesses, chart a year of progress in visual capabilities, and publicly release ZeroBench at https://zerobench.github.io/.
Lay Summary
AI systems that can interpret images are improving rapidly, but they can still make simple visual mistakes, such as miscounting objects. At the same time, many popular image-based tests are now close to being solved, making it harder to tell whether new systems are genuinely better. We introduce ZeroBench, a test of 100 visual reasoning questions, built to be extremely challenging for today’s strongest AI systems. Each question was manually created, reviewed, filtered for difficulty, and later checked by other researchers to reduce errors. When ZeroBench was released, all 20 evaluated systems failed to answer any question correctly on their first attempt, showing that the test was beyond contemporary model capabilities. We then tracked 26 newer systems over the following year and found clear progress, although the best systems still solved only a minority of questions. Most failures came from visual interpretation: counting objects and recognising details. ZeroBench provides a yardstick for measuring future progress in visual AI systems and understanding where they still fall short.