Position: Alignment Needs Rule-Class Routing Before Preference Learning
Abstract
This position paper argues that alignment pipelines should classify rules before they aggregate preferences. RLHF, Constitutional AI, DPO, RLAIF, and Deliberative Alignment differ technically, but each tends to turn contested normative input into one global policy. Social-choice theory explains why this is unsafe on unrestricted domains: aggregation can help when there is public convergence, but it cannot decide which questions are eligible for aggregation. Conitzer et al. reframe alignment as social choice; we agree, and argue social choice identifies a boundary, not a solution. Arrow, Gibbard–Satterthwaite, and Sen establish that aggregation over unrestricted value disagreement has no procedure satisfying minimal democratic, strategic, and liberty-preserving conditions. The eligibility decision is institutional: the system designer must state who is authorized to decide whether a rule is aggregable, user-configurable, or non-negotiable. We propose a routing layer for behavioral rules. Class I rules are public prohibitions suitable for aggregation; Class II rules concern reasonable disagreement and should support configurable defaults; Class III rules protect rights or vulnerable users and should be implemented as constraints with appeal mechanisms. The paper connects impossibility results to alignment design, compares current methods through this lens, and gives a six-step routing procedure with a five-item disclosure checklist.