Covert Trait Propagation Is Representation Alignment: Mechanistic Evidence from Hidden-Channel Distillation
Kargi Chauhan ⋅ Aditya Shah
Abstract
A student model trained on pure uniform noise can still inherit its teacher's digit-classification ability, provided the two share initialization. Previous work proves this transfer is guaranteed when the teacher's learning rate is small enough, but does not explain where in the network the channel lives or what sets its capacity. Working in an MLP distillation setting on MNIST, we show these channels are not purely informational: geometric alignment gates access to the information the channel carries. Shared initialization makes the output projection $W_2$ a common coordinate key, and KL gradients reshape the student's input projection $W_0$ until its hidden representations align with the teacher's. We call this covert trait propagation (CTP). Five experiments support this mechanism: channel closure tracks weight drift, not teacher accuracy; freezing $W_0$ destroys transfer while freezing $W_2$ leaves it intact; multi-teacher ensembles cancel out despite each teacher carrying comparable label information; and CKA tracks student accuracy at $r{=}0.98$ across a continuous initialization sweep. Applying the same geometric lens to cross-token behavioral entanglement (CTBE) in instruction-tuned LLMs, we find the effect is activated by alignment training, acting on an inherited substrate, and that the standard log-ratio metric produces an apparent frequency bias that is largely a circularity artifact.
Chat is not available.
Successful Page Load