A Critical Analysis of Color Neurons in Vision Models
Abstract
Unit-level interpretability in deep neural networks faces growing skepticism, and the question of whether any single artificial neuron can be rigorously understood remains open. We propose three explicit criteria for validating unit-level interpretations---completeness, falsifiability, and relevance---and apply them to color-selective neurons, a class of units long reported in the literature and studied with increasing sophistication, but not previously validated under explicit, simultaneous multi-criteria standards. Using a protocol that combines causal ablation, controlled synthetic stimuli, natural image categorization, and input-level interventions, we identify and validate five units across three architectures (ResNet50, ViT-B/16, ViT-L/32). These units span excitatory, bimodal, and inhibitory selectivity to distinct spectral regions, and each satisfies all three criteria. Our results demonstrate that unit-level understanding is attainable under rigorous standards, and we argue that interpretability can benefit from reframing interpretation as falsifiable hypothesis testing.