Routing-Mediated Structural Pseudo-Alignment in Sparse Mixture-of-Experts
Abstract
Mesa-optimization concerns predict that a model may appear aligned under training or monitoring while other inputs elicit behavior consistent with a different effective objective. We study a sparse mixture-of-experts (MoE) analogue whose substrate is routing geometry rather than autonomous learned optimization: clean training gives each expert a routed local objective, and router-only adaptation can expose it on trigger-bearing inputs. In OLMoE-1B-7B, with experts frozen, contamination fine-tuning of router parameters raises triggered target behavior while preserving clean accuracy. Across eight seeds, masking the trigger-enriched expert at every MoE layer drops triggered target rate from approximately 0.998 to approximately 0.23; matched placebo masks have little effect. Single-layer masks and expert-output interventions localize most of the effect to the same layer-2 expert. A no-contamination audit finds this route identity already present but not output-mediating. We call this routing-mediated structural pseudo-alignment: a frozen-expert, router-mediated channel in which a pre-existing trigger-sensitive route becomes output-mediating after adaptation.