Sparse Groundings: A Benchmark for Auditing Visual Representation Claims
Abstract
Modern multimodal systems rely on internal visual representations, but claims about what those representations contain are typically supported by probes or interventions read from the patches covering an object. These readouts can conflate object evidence with signals that are easier to miss: information broadcast across the image, background cues, patch coordinates, or the evaluator's own pooling rule. We introduce Sparse Groundings, an audit benchmark for visual representation claims built around three controlled readouts at the same patch count. The standard object readout is compared with a blank same-footprint readout that preserves selected patch coordinates while removing image content, and a matched-background readout that preserves image context and patch count while removing the object. Each result is recorded as a claim card summarizing the tested property, controls, method, and verdict, and aggregated into an evidence profile rather than a single groundedness score. We instantiate the benchmark on controlled tests for identity, colour, position, scale, spatial relations, counting, and distractor robustness across six vision encoders and two sparse-autoencoder bases. The controls change conclusions in three places. In one representative case, a SigLIP 2 probe predicts which cell of a 7 by 7 grid contains the object with 0.991 balanced accuracy from object features, while a uniform blank image read from the same patch locations reaches 0.998 — a pattern that holds at ≥ 0.978 across all eight representations. Identity and attribute content is recoverable from non-object patches as well as from object patches on every encoder, and structure-family claims partition encoders on a different axis from position-family claims. A natural-image bridge readout on LVIS object boxes recovers the same encoder partition — five of six encoders show object-pool features above the strongest control, with MAE the same matched-background-dominated outlier — so the regimes are not artefacts of the grey-canvas protocol. Task-matched evidence profiles align with ADE20K segmentation and NYUv2 depth, do not predict iNat fine-grained classification, are fragile on FSC-147 counting, and offer evidence for RefCOCO+ grounding under the current pooled-feature protocol. Sparse Groundings therefore audits not whether a model is grounded in general, but which visual claim survives which controls, by which method, and for which downstream use. As multimodal systems become more central to perception-heavy applications, understanding what they perceive about the world will require interpretability tools that are precise about the claim being tested and robust to the controls that could explain it away.