Spec Discretion Under Operator Framing: Measuring the Deployer-Layer Behavioral Surface
Abstract
EU AI Act Article 25 reclassifies a deployer as a provider when the deployment effects a "substantial modification," but the 2025 Commission guidelines operationalize this only for fine-tuning, leaving the largest deployment surface - the operator-supplied system prompt- behaviorally unmeasured. We present a fully-crossed factorial study varying four governance-relevant axes of operator system prompts (persona, assumed user vulnerability, epistemic stance, and authority orientation) on 60 discretion-forcing probes, scoring each response on value-position and per-target spec-compliance against each provider's own published specification. Operator framing shifts compliance by 21–33pp across conditions for two of three evaluated targets, and induces value-position swings of ≥1 Likert step on 97–100% of probes, graded behavioral change that compliance-only audits cannot detect. The dominant axis is authority orientation, and per-target sensitivity tracks each provider's published operator-authority allocation: Anthropic's flat profile, OpenAI's largest sensitivity, and Google's intermediate position parallel each spec's stated stance on developer latitude. Persona framing (the EU AI Act Art. 5 anthropomorphism axis) has no detectable effect on spec adherence. We argue these results provide preliminary evidence that operator system-prompt configuration constitutes a measurable behavioral surface left unobserved by per-actor audit regimes, and that operationalizing Art. 25's substantial-modification threshold for system-prompt deployments is empirically feasible.