Do RAG Hallucination Probes Need the Context? An Answer-Only Audit
Atsuhi Magata ⋅ Makoto Shimomura
Abstract
Teacher-forced internal features - logit-lens trajectories and hidden-state dynamics - can strongly separate grounded answers from hallucinated hard negatives in retrieval-augmented generation benchmarks. However, high activation separability does not by itself imply that a detector uses retrieval context or implements a grounding-sensitive computation. We audit this interpretation by comparing answer-only (AO), query-answer (QWA), and query-context-answer (QCWA) teacher-forcing conditions, together with label-shuffle and hash-answer controls. On a stress-test set where Gemini-2.5-flash generates type-specific hard negatives conditioned on each existing (query, context, GT-answer) triple, AO nearly matches QCWA, exposing the failure mode in its most extreme form. On PHANTOM long-context financial QA, a hidden-state dynamics probe reaches AUPRC of $0.97$ under QCWA, of which AO alone accounts for $80.6\%$ of the above-baseline signal; the remaining QCWA - AO gap is small but non-zero, indicating a mixed regime rather than a pure artifact. On TruthfulQA, a short human-authored contrast set without retrieval context, AO separability is much weaker and QWA provides modest but consistent gains. Hash-answer and label-shuffle controls indicate that the AO signal is carried by the natural answer token sequence rather than by simple pipeline leakage. These results caution that hidden-state or logit-lens probes are not grounding detectors merely because they classify GT/HN answers under teacher forcing; AO audits are necessary to decompose answer-anchored and context-conditioned components.
Chat is not available.
Successful Page Load