Revisiting the Platonic Representation Hypothesis: An Aristotelian View
Abstract
The Platonic Representation Hypothesis suggests that representations from neural networks are converging to a common statistical model of reality. We show that the existing metrics used to measure representational similarity are confounded by network scale: increasing model depth or width can systematically inflate representational similarity scores. To correct these effects, we introduce a permutation-based null-calibration framework that transforms any representational similarity metric into a calibrated score with statistical guarantees. We revisit the Platonic Representation Hypothesis with our calibration framework, which reveals a nuanced picture: the apparent convergence reported by global spectral measures largely disappears after calibration, while local neighborhood similarity, but not local distances, retains significant agreement across different modalities. Based on these findings, we propose the Aristotelian Representation Hypothesis: representations in neural networks are converging to shared local neighborhood relationships.
Lay Summary
As AI models grow, research has suggested that they may learn similar "views of the world". This paper shows that some of this apparent similarity can be caused by model size rather than true shared understanding. After correcting for these effects, this paper finds that models do not necessarily share the same global representation, but they often agree on which examples are similar to each other.