How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study
Abstract
While AI-generated hallucinations pose considerable risks, the underlying cognitive mechanisms by which humans can successfully recognize or be misled by these hallucinations remain unclear. To address this problem, this paper explores humans' neural dynamics to characterize how the brain processes hallucinated content. We record EEG signals from 27 participants while they are performing a verification task to judge the correctness of image descriptions generated by a multi-modal large language model (MLLM). Based on an averaged event-related potential (ERP) study, we reveal that multiple cognitive processes, e.g., semantic integration, inferential processing, memory retrieval, and cognitive load, exhibit distinct patterns when humans process hallucinated versus non-hallucinated content. Notably, neural responses to hallucinations that were misjudged versus correctly judged by human participants showed significant differences. This indicates that misjudged AI-generated hallucinations failed to trigger the standard neurocognitive fact verification pathway. The detailed code can be accessed openly through https://github.com/Promise-Z5Q2SQ/EEG-Hallucination.
Lay Summary
This paper studies how people detect mistakes made by AI systems that describe images. Although modern AI can generate very realistic image captions, it can also produce false or misleading statements, i.e., “hallucinations”. We wanted to understand what happens in the human brain when people successfully notice these errors, and why they sometimes fail to do so. To investigate this, we recorded brain activity from 27 participants while they checked whether AI-generated image descriptions were correct. We found that the brain responds differently when processing truthful versus hallucinated descriptions. In particular, recognizing hallucinations involved stronger signals related to understanding meaning, recalling knowledge, and making decisions. However, when participants were fooled by hallucinated content, these neural responses were much weaker or absent. Our findings suggest that successfully detecting AI misinformation depends on a specific chain of cognitive processes in the brain. When this process is not fully activated, people may accept incorrect AI-generated information as true. This work helps improve our understanding of how humans interact with AI systems and may contribute to building safer and more trustworthy AI tools in the future.