Spectral Stratification of Semantic Abstraction in Vision-Language Models
Abstract
Vision-language models now serve as general-purpose semantic embedding spaces, but how conceptual abstraction is organized within such spaces remains poorly understood. Here, we show that abstraction in these embeddings is spectrally stratified. Specifically, low-rank principal-component subspaces encode broad conceptual structure, while higher-rank components carry progressively finer-grained distinctions. This pattern holds across multiple contrastive vision-language models and taxonomic datasets. Consistent with this stratification, rank-based modulation produces predictable behavioral shifts: retaining only the leading components preserves coarse retrieval but degrades fine-grained retrieval, while removing them selectively disrupts abstract category structure. Furthermore, rank-selective routing improves retrieval when the selected subspace matches the target level of abstraction. Notably, compact spectral subspaces, which preserve taxonomic structure while truncating high-rank residuals, exhibit substantially better alignment with human abstraction behavior than the full-rank model. Together, these results reveal that vision-language embeddings are not semantically homogeneous. Instead, hierarchical abstraction is stratified along the spectral geometry, and the structure most aligned with human cognition is concentrated in a compact subspace, suggesting that human-aligned semantics are a recoverable substructure of, rather than a property of, the full embedding.