What makes the whole? Probing Attribute-Level Compositionality in LLM Judges
Savita Bhat ⋅ Vasudeva Varma
Abstract
LLM judges produce a multidimensional evaluation in the form of an overall quality score, along with fine-grained attribute-level judgments, implicitly hinting at compositionality. We investigate whether LLM judges genuinely compose part-evaluations into whole, or instead produce decomposable-looking evaluations through a different mechanism. Across diverse multi-dimensional evaluation settings (summarization, dialogue, and essay), we use a four-test diagnostic framework: 1) to probe whether overall scores can be reconstructed from a combination of attribute-level scores, 2) whether the inferred composition is stable and human-aligned, 3)whether attribute-level scores provide genuine diagnostic value, and 4) to evaluate sensitivity towards targeted trait degradation. Across summarization, dialogue, and essay evaluation (4 models × 3 prompts × 4 datasets), we find that LLM judges produce highly decomposable outputs ($R^2$ = 0.755-0.934), yet their inferred composition diverges sharply from humans (Kendall's $\tau$ between LLM and human trait orderings: -0.07 to 0.20). Counterfactual perturbations for targeted trait degradation establish that decomposability does not imply causal sensitivity. The dominant factor is not model capacity, but the geometry of the rubric. The externally-imposed compositional structure explains 17× more variance in compositional behavior than model choice ($\eta ^ 2 $= 0.384 vs. 0.022). The decomposability of structured output is a weak indicator of genuine compositional reasoning, and the inductive biases imposed by the task structure can dominate the model-side capacity.
Chat is not available.
Successful Page Load