Stop Reporting System-Level AI Reasoning as Individual Model Capability
Abstract
GPT-5.5 scores 41.4% on Humanity's Last Exam. GPT-5.5 Pro, the same model with parallel test-time compute and tool access, scores 57.2%. Both appear on leaderboards as "GPT-5.5." That 15.8-point gap is not a model improvement; it is a system-composition difference wearing a model-capability label. This paper argues that frontier reasoning scores are produced by assemblies of models, test-time compute, tools, and verifiers, yet the field attributes them to a single model name. We develop a four-level hierarchy of inference complexity, present a pilot audit of 26 entries across 11 benchmarks revealing pervasive undisclosed system-level composition, ground the argument in distributed cognition theory, and propose a compute-normalized reporting protocol (SCRP) as a minimum standard for scientific credibility.