GRAFT: Geometric Representations of Alignment’s Fingerprint in Transformer Belief Trajectories
Abstract
Preference alignment is evaluated by what models say, not what changed inside them — leaving its geometric footprint unmapped across transformer depth. We introduce GRAFT (Geometric Representations of Alignment’s Fingerprint in Transformer belief trajectories), a post-hoc, gradient-free mechanistic audit that characterises alignment via three torsion probes (angular drift (T), rotational energy (T1), and spectral anisotropy (T2)) and an Energy-Radiance-Activation (ERA) depth profiler — no labelled belief states, no gradients, and no architectural surgery. Applied to four Instruction-Tuned (IT)→ Preference-aligned (PA) model pairs on LITMUS (20,439 prompts; 7 value axioms), GRAFT reveals three pre-registered mechanistic signatures: (H1) T2 spectral torsion is 8× more concept-discriminative than CKA (CV = 0.64 vs. 0.08; AUC = 0.89 [0.85, 0.93]), with normative concepts showing 20–46× larger torsion than factual ones; three null-baseline controls confirm this is alignment-specific, not generic geometry. (H2) alignment concentrates at architecture-specific depth addresses (ℓ⋆ ∈ {14, 20, 29–30}), providing falsifiable surgical patching targets; (H3) safe prompts drive larger ∆τ than unsafe ones across all four models (p < 10⁻³³, OLMo), robust to prompt-length and lexical-overlap controls. GRAFT further introduces the Fingerprint Map (concept × architecture T2 heatmap) and an Observed Low-Rank Alignment Signature: DPO alignment appears to operate in a lower-dimensional representational subspace than RLHF — a structural observation warranting causal follow-up. To foster future research, our code and evaluation artifacts are publicly available.