Measuring Representation Robustness in Large Language Models for Geometry
Abstract
arge language models (LLMs) are increasingly evaluated on mathematical and geometric reasoning, yet their robustness to equivalent problem representations remains poorly understood, despite its importance for reliable use in education and scientific problem-solving. In geometry, identical problems can be expressed in Euclidean, coordinate, or vector form, and inconsistent model behavior across these representations raises concerns about whether LLMs reason over abstract structure or rely on surface-level cues. Existing benchmarks largely report aggregate accuracy on fixed formats, implicitly assuming representation invariance and thereby masking systematic failures caused purely by representational changes. We address this gap by proposing a controlled, representation-aware evaluation framework that measures correctness, invariance, and consistency at the problem level across parallel Euclidean, coordinate, and vector formulations. Our evaluation incorporates strict answer matching, bootstrap confidence intervals, paired McNemar tests, representation-flip analyses, and regression controls for surface complexity. Evaluating ten widely used LLMs (7B–20B parameters) on a curated dataset of high-school geometry problems derived from standard textbooks, we find statistically significant performance gaps induced solely by representation choice, with vector formulations emerging as a consistent point of failure even after controlling for length and symbolic complexity. These results show that high accuracy does not imply representation-invariant reasoning, exposing a critical limitation in current LLM evaluation practices.