Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports
Phongsakon Mark Konrad ⋅ Toygar Tanyel ⋅ Serkan Ayvaz
Abstract
Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demonstration environment changes. Across three open instruction models, wrong-source demonstrations move reports toward the source answer family, while explicit mechanism binding reduces this pull. Self-report benchmarks should include environment-shift invariance tests under fixed intervention before treating accuracy as evidence for an autonomous report mechanism.
Chat is not available.
Successful Page Load