Beyond Expert Benchmarks: Decomposing Query Phrasing Effects on Legal AI Performance
Abstract
Legal AI benchmarks typically evaluate models on expert-phrased legal prompts, while Access-to-Justice users often describe legal problems in ordinary or emotionally distressed language. We construct controlled triplets from LegalBench: the original \textit{Expert} prompt, a \textit{Naive-Calm} rewrite, and a \textit{Naive-Distressed} rewrite generated from the calm version. We preserve LegalBench's original prompt templates and scoring, changing only the inserted scenario text. Evaluating three frontier models, we find that phrasing effects are not only domain-dependent but also model-dependent: models exhibit qualitatively distinct sensitivity profiles, from near-complete robustness to significant monotonic degradation under lay phrasing. Error analysis reveals four distinct flip patterns---including cases where distressed phrasing accidentally improves performance by adding clarifying emphasis---suggesting that phrasing effects operate through identifiable mechanisms rather than uniform noise. These results suggest that expert legal phrasing is not a neutral benchmark baseline and that legal AI evaluation should explicitly test robustness across user-language registers.