Bargaining Over Bias: Cooperative Contracts for AI Fairness
Abstract
Fairness-critical AI systems require multi-stakeholder agreement on parameters such as disparity thresholds and harm weights, yet these are almost always set unilaterally by system designers. We formalise this as a mechanism design problem and introduce the Fairness Debt Contract (FDC): a negotiable tuple ⟨δ, H, α⟩ encoding a fairness threshold, harm weight, and accuracy budget. We compare four negotiation protocols and propose a hybrid Median-then-Nash (MtN) protocol. Agent-based simulations (150 rounds, 20 stakeholders) show that MtN produces approximately 60% higher (1.6×) social welfare (defined here as the sum of stakeholder preference satisfaction at the negotiating table, distinct from downstream demographic outcomes) than Nash bargaining (4.06 vs. 2.54) and converges in every round, whereas unanimity never converges. MtN's first stage, median scoping, is verifiable and auditable, making it a candidate governance primitive for Technical AI Governance (verifiability, operationalisation) and Cooperative AI (bargaining, trust).