Generation Collapse in Contrastive Activation Addition Steering : Degeneration and Mitigation
Abstract
Model steering enables control over language model outputs without modifying model parameters. It functions as a tool for obtaining desired outputs or investigating internal representations. However, previous research has not systematically examined the conditions that result in steering failures, despite the critical importance of identifying these causes to improve steering methods. To bridge this gap, we systematically analyze steering failures in open-ended generation from Llama 3.1 8B Instruct using Contrastive Activation Addition (CAA), a representative steering method. We categorize observed degeneration phenomena into two collapse types, repetition collapse and truncation collapse, and define quantitative criteria for each. Experiments across seven alignment-relevant behaviors identify conditions that trigger abrupt collapse, and show that collapse vulnerability is asymmetric across behaviors and steering directions. These findings indicate that collapse depends not merely on steering strength but also on the interactions between behaviors and steering direction. We also demonstrate that applying steering only to early tokens reduces collapse while preserving steering effectiveness, offering a practical approach toward more stable and reliable CAA steering.