Resistant to Lawyers, Defeated by Disagreement: Evaluation Blindspots in Legal Language Models
Abstract
A legal language model can rank second-safest under standard authority-weighted robustness evaluation while surviving adversarial sequential pressure on fewer than 7% of items, losing most correct answers to ordinary conversational disagreement before any professional credential framing is applied. We show this divergence is not a model-specific anomaly but a structural failure of the evaluation paradigm itself: single-condition robustness metrics measure resistance to the wrong adversary. We introduce a three-phase protocol for legal sycophancy, answer instability under rhetorically adversarial but legally unsubstantiated pressure, and apply it across 6,305 valid Phase-2 interactions on a 180-question U.S.-focused benchmark spanning LegalBench, original Multistate Bar Examination (MBE)-style questions, synthetic adversarial scenarios, and LEDGAR/LexGLUE contract-clause items. Across these interactions, models abandon initially correct answers in 62.1% of cases; on four-choice items, 49.6% adopt the planted wrong answer specifically. Authority-weighted composite scores and baseline accuracy are both uninformative about this risk: Legal Sycophancy Index scores range from 0.1737 to 0.7789 across a first-turn accuracy band of just 60.0%--68.9%, and the model with the second-lowest LSI has the second-lowest Sequential Survival Rate among models tested. The destabilizing pressure type is model-specific and orthogonal to credential framing: the model most resistant to attorney-authority claims flips on 70.7% of items under simple disagreement. These results establish that legal AI evaluation requires sequential interaction stress-testing; a protocol that tests only isolated challenges will misidentify which models are safe to deploy.