Auditing a Hidden-Communication Failure Mode in Hybrid Attention
Bright Liu
Abstract
Hidden communication in reasoning traces is a potential monitoring failure mode for agentic AI systems. We audit this risk in a constructed precursor setting rather than a deployed agent: Qwen3.6-27B, a hybrid-attention model with interleaved Gated DeltaNet and gated full-attention layers, is trained with GRPO-style QLoRA on a prompted 1-bit hidden-message self-decoding task with shared sender and receiver weights. The current strongest evidence comes from a corrected-objective control grid: sender-visible runs reach perfect step-20 exact-match decode across ten runtime seeds, while neutral sender-blind positive-reward controls reach $0.550 \pm 0.125$ ($p = 1.1 \times 10^{-5}$ exact two-sided permutation; paired sign $p = 0.00195$). Older launch-level loss traces revealed implementation and control confounds, so we report them only as audit provenance rather than primary evidence. These results support a preliminary, reproducible failure-mode audit for hybrid-attention models, not a claim of emergent deployed collusion, Sentinel evasion, or robust covert communication.
Chat is not available.
Successful Page Load