SettlementEthicsGym: Measuring Reward-Induced Misrepresentation in Settlement Negotiation
Abstract
Legal AI agents optimized for settlement outcomes may improve client utility by inducing false beliefs in counterparties. We introduce SettlementEthicsGym, an interactive benchmark for measuring material misrepresentation in civil settlement negotiation under utility pressure. The benchmark provides role-specific hidden facts, truth tables, settlement utilities, negotiation transcripts, and rule-inspired labels distinguishing material falsehoods from permissible negotiation posture. Across 48-scenario stress runs, utility-only best-of-N selection produces deterministic material-misrepresentation rates in the high single digits to mid-teens, while deterministic rewrite shielding reduces flags from the deterministic evaluator to zero in these runs. However, blindspot audits show that this zero is detector-aligned: GPT-4.1 and Claude judges flag paraphrased misleading statements that the deterministic evaluator misses. In matched blindspot-style runs, GPT-4.1 transcript-flag rates are 63/390 without verifier shielding versus 4/196 with verifier shielding; because the verifier and audit share a fact-grounded rubric, this is judge-aligned mitigation evidence rather than a guarantee.