Compositional Underdetermination in AI Agents: When Behavioral Success Is Not Compositional Evidence
Abstract
Modern agent evaluations often report end-to-end success rates, yet such scores rarely establish whether an agent has acquired reusable compositional structure or brittle task-specific routines. We name this gap compositional underdetermination: behavioral success on compositional tasks can remain compatible with multiple incompatible explanations of how an agent plans, calls tools, retrieves information, represents state, or composes safety constraints. We formalize the gap as a relation on behavioral equivalence classes and argue that compositionality claims for agents require paired behavioral and diagnostic evidence. The paper contributes an agent-specific taxonomy of compositionality surfaces, a minimum reviewable diagnostic for each surface, and a reporting standard that existing benchmarks can adopt. The result is a lightweight reporting protocol for agent benchmarks that connects compositional generalization, interpretability diagnostics, and safety-constraint evaluation without requiring a new benchmark.