Reusable Uncertainty Representations Support Metacognitive Behavior in Llama-3.3-70B
Abstract
Knowing what you do not know is a hallmark feature of intelligent behavior. This is grounded in uncertainty monitoring, a core component of metacognition that guides decisions such as seeking help or withholding unreliable responses. For Large Language Models (think chatbots), this question is increasingly important, as these systems are deployed in consequential settings. Yet it remains unclear whether confidence-related behaviors in LLMs reflect reusable internal uncertainty representations or task-specific heuristics shaped by prompts and surface cues. We study this question in Llama-3.3-70B-Instruct by identifying residual-stream directions associated with answer-logit uncertainty during factual multiple-choice question answering, then testing whether the same directions transfer to metacognitive readouts, including stated confidence and delegation. We find evidence for a reusable uncertainty-related representation: the directions transfer across datasets and readouts, acquire abstention-related output semantics, and can be steered to shift delegation behavior. Ablation provides weaker and more selective evidence, suggesting that these directions are sufficient to influence delegation but not uniquely necessary. Together, these results support a constrained account in which LLMs can reuse internal uncertainty- related representations across metacognitive contexts, while the behavioral expression of those representations depends on the task.