Tracing Psychometric Inference in Large Language Models
Chih-Hao Hsu ⋅ Feng-Chun B Chou ⋅ Pin-Hao A Chen
Abstract
Do large language models represent psychological constructs internally, or do they only generate outputs that resemble psychometric structure? We address this question using a cross-persona paradigm across 14 models (Llama3 and Qwen2.5, 0.5B--14B, base and instruct). Given responses on one psychological scale from real individuals ($N = 272$), models predict responses on six other scales. At the behavioral level, LLMs reproduce human cross-scale correlation structure and systematically amplify it, with model-generated correlations exceeding human estimates even after correcting for measurement attenuation. We then examine whether this structure is reflected internally. Contrastive direction analysis reveals an organized geometry in activation space aligned with psychometric relationships. This structure emerges in large instruct models but is not observed in base models. Across models, geometry strength predicts behavioral amplification ($r = 0.68$, $p = 0.008$; partial $r = 0.69$ controlling for log size), independent of model size. To relate internal structure to output behavior, a matched activation-probing paradigm shows that representational amplification is less variable than behavioral amplification ($1.38$--$1.77$ vs $0.53$--$1.53$). A synthetic control with known ground-truth structure shows that this range survives subtraction of a ridge-probe baseline ($\approx 1.3$), with adjusted slopes still predicting behavior at $r = 0.88$. We term the systematic representation-to-behavior gap \emph{readout attenuation}. Together, these findings suggest that LLMs encode structured representations aligned with psychological constructs, while differences in output primarily reflect how these representations are read out.
Chat is not available.
Successful Page Load