The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Abstract
Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral tendencies resembling human traits like agreeableness and self-regulation. Understanding these patterns is crucial, yet prior work primarily relied on simplified self-reports and heuristic prompting, with little behavioral validation. In this study, we systematically characterize LLM personality across three dimensions: (1) the dynamic emergence and evolution of trait profiles throughout training stages; (2) the predictive validity of self-reported traits in behavioral tasks; and (3) the impact of targeted interventions, such as persona injection, on both self-reports and behavior. Our findings reveal that instructional alignment (e.g., RLHF, instruction tuning) significantly stabilizes trait expression and strengthens trait correlations in ways that mirror human data. However, these self-reported traits do not reliably predict behavior, and observed associations often diverge from human patterns. While persona injection successfully steers self-reports in the intended direction, it exerts little or inconsistent effect on actual behavior. By distinguishing surface-level trait expression from behavioral consistency, our findings challenge assumptions about LLM personality and underscore the need for deeper evaluation in alignment and interpretability.
Lay Summary
In humans, personality is typically measured by asking people to fill out self-report questionnaires about themselves, and those answers help predict how they will behave in real-world tasks. Recent work shows that modern Large Language Models (LLMs), such as ChatGPT, also give distinct personality answers on these same questionnaires. But do those self-reports actually predict how the LLM behaves, and do they follow the same patterns as in humans? We tested 18 modern Large Language Models (LLMs) by combining personality questionnaires with realistic tasks drawn from psychology. We found: 1) Training makes an LLM's self-described personality more stable, more internally consistent, and more socially desirable. 2) Yet these self-reports rarely predict actual behavior, and when they do, the direction often disagrees with what human psychology would predict. 3) Prompts that try to give the LLM a personality shift what it says about itself, but have much weaker effect on what it actually does. We call this gap the personality illusion: today's LLMs sound like they have a coherent personality, but those words are not a reliable guide to their behavior.