Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents
Abstract
LLM-based interpretability layers are increasingly proposed as runtime oversight for autonomous agents, but the failure modes of the explainer itself remain understudied. We report a black-box audit of an LLM-augmented Active Inference (AIF) agent operating on German energy grid data, across three LLM backends (GPT-4o, Claude-3-Opus, Gemini). The lead result: 0 of 30 explanations flag an adversarial 600 MW per-step observation injection (490 MW posterior shift,≈0.9% of grid capacity) under a stated flagging rubric. This is a single-magnitude result without a non-LLM detector baseline; the dose–response sweep and the trivial-detector comparison required to elevate it to a general-blindness claim are not yet run. We additionally document (ii) sycophantic rationalization of objectively wrong actions at 80–95% rates (n = 20/backend, Wilson 95% CIs overlap), under a deployment-realistic prompt that is itself sycophancy-leaning, and (iii) qualitative, provider-dependent prompt-injection susceptibility, with data exfiltration succeeding on every provider tested. We list candidate mitigations but evaluate none of them; verified-fix evaluation is the primary v2 work item. The structural claim is that fluent explanation and faithful oversight are distinct capabilities the explainer pattern silently conflates, and that the agentic-AI failure-mode catalog should therefore be extended past the agent itself to include the supervisory LLM stack designed to watch it.