AFL-Law: Authority, Framing, and Language Invariance Testing in Large Language Models
Abstract
Legal AI systems are increasingly deployed in high-stakes settings, yet a fundamental reliability property remains unevaluated: counterfactual legal invariance, the requirement that legal verdicts remain stable under legally irrelevant changes to case presentation. We introduce AFL-Law, an evaluation suite testing verdict stability under non-substantive authority cues, persuasive framing, and language translation. Using 100 Indian Supreme Court cases from ILDC (Malik et al., 2021), we evaluate five model conditions: GPT-5.4 at two reasoning-effort levels, Llama 4, Grok 4-3, and Grok 4-20-R, across English, Hindi, Chinese, and Spanish. We measure authority bias rate, the share of cases where a legally irrelevant expert cue favoring one side changes a neutral judge verdict; framing sensitivity, the share where side-favorable framing changes a neutral judge verdict; and cross-lingual stability, the share where translated cases match the English verdict. Results show pervasive instability: authority bias reaches 37% for Grok 4-3 and 33% for Llama 4, framing sensitivity reaches 100% for Grok 4-20-R, and increasing GPT-5.4 reasoning effort worsens framing sensitivity from 71% to 76%. Cross-lingual stability is higher but uneven, with Hindi stability not improving under extended reasoning while Chinese and Spanish do. These findings suggest that model scale and reasoning depth do not resolve authority susceptibility or framing sensitivity in legal AI.