How Do LLMs Distinguish Normative Ethical Frameworks Internally?
Abstract
Aligned LLMs apply different ethical frameworks (deontology, utilitarianism, virtue) when making moral judgements; how this content is organised in hidden state determines whether framework-specific failures surface diagnostically or hide in a broader signal. Two views bracket the question: distinct frameworks in separate subspaces, or all collapsing onto a single endorsement axis. We find a third structure across 13 open-weight models spanning 7–72B and six moral systems: the causally-used moral signal lies in a shared low-dimensional endorsement subspace where each system holds its own direction (geometrically distinguishable) yet directions substitute causally for one another with target-specific asymmetry (functionally substitutable). The subspace is recoverable in pretrained base models and reshaped at supervised fine-tuning rather than later preference-learning. Three implications follow: defences anchored to a single moral direction miss orthogonal-basis attacks within the same subspace; "moral steering" reshapes endorsement and framework emphasis at comparable rates while preserving multi-framework reasoning; and direction-level alignment leverage sits upstream of preference learning at supervised-fine-tuning data construction.