Paper #78: Creative Collision: Directorial Persona Steering and Competition in Large Language Models
Abstract
Activation steering enables large language models to be controlled at inference time by injecting semantic directions into internal activations, yet prior work has largely focused on isolated steering vectors. We study a richer regime in which two semantically opposing steering directions are superimposed, a setting we call Creative Collision. Using curated screenplay-derived corpora, we construct directorial persona vectors for Steven Spielberg and Martin Scorsese through mean-difference activation contrast, representing optimistic, redemptive versus dark, morally ambiguous narrative tendencies. We interpolate between these vectors and apply them to the residual stream of a 14B parameter decoder-only transformer during generation. Across evaluations of moral valence, coherence, stylistic structure, directional dominance, and vector geometry, we report three principal findings. First, the Spielberg vector exhibits strong directional dominance, suppressing Scorsese influence across most of the interpolation range. Second, intermediate collision points improve coherence relative to pure single-vector steering at high steering magnitudes, consistent with a norm-reduction effect arising from non-antipodal vector geometry. Third, both personas localise maximally to the same upper-middle transformer layer, suggesting a shared moral-tone processing substrate. We further observe that partial vector collision produces morally richer generations than either pure persona extreme, indicating that competing semantic directions can amplify narrative complexity rather than merely interpolate between styles. These results provide new insight into how competing high-level semantic representations interact within transformer residual streams and suggest new directions for controllable creative generation and mechanistic interpretability.