Inoculation Prompting Composes With Alignment Pretraining
Abstract
Alignment interventions can be applied during both pretraining and post-training to steer safety-relevant properties of models. Two such interventions are alignment-pretraining, which curates pretraining data to instil aligned priors in base models, and inoculation prompting, which elicits unwanted behaviours during fine-tuning to paradoxically suppress them at test time. Whether they compose is untested, but matters as interference between alignment techniques could erode their utility. We study their composability by applying inoculation prompting to two 6.9B alignment-pretrained models and a 6.9B unfiltered baseline. The interventions compose without measurable interference: equivalence testing (TOST) shows inoculation prompting's effect on alignment-pretrained models matches its effect on the unfiltered models within 3 percentage points across all three trait-eliciting prompt strengths, with strong inoculation reducing emergent misalignment across all three base models from 6.6-7.3% (per-base Wilson 95% CIs all within [5.3, 9.0]) to 0.2-0.3% (95% CIs all within [0.1, 0.7]). We further show that trait-eliciting inoculation prompts reduce emergent misalignment beyond placebo prompts and that this protection partially persists even when the inoculation prompt is present at test-time. These results suggest that inoculation strengthens aligned models against emergent misalignment through mechanisms beyond conditionalisation alone.