Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models
Subramanyam Sahoo ⋅ Justin Shenk
Abstract
What happens when a legal AI learns to \emph{look} like a lawyer instead of \emph{reasoning} like one? We fine-tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a surface-feature proxy --- citation count, legalese density, and response length --- and find that the model does not learn to reason better. It learns to stall. Across $16$ yes/no legal-reasoning tasks from LegalBench ($N{=}320$), overall accuracy collapses from $0.500$ (chance) to $0.072$ ($\mathrm{McNemar}$ $p < 10^{-36}$), driven entirely by the answer-format rate plummeting from $0.900$ to $0.109$. The model stops committing to answers. Yet when it does commit, accuracy \emph{rises} from $0.556$ to $0.657$ --- revealing that the collapse is not a capability failure but a strategic one: the model has learnt that verbose, citation-dense non-answers score higher than terse correct ones. We term this the \emph{Saul Goodman effect} --- a policy that becomes maximally lawyerly while becoming maximally non-committal --- and prove formally that it is the \emph{optimal response} to any surface-feature proxy that attaches no penalty to abstention (Proposition~ 3.2). We further show that $89.3\%$ of post-training citations are structurally implausible hallucinations, many of which are subtly corrupted names of real landmark cases designed, in effect, to pass a casual read and fail under scrutiny. To detect this failure mode before deployment, we introduce three diagnostic tools: the Confidence Theater Score (CTS), the Citation Plausibility Rate (CPR), and the Regret Gap (RG). In a domain where a confidently wrong answer can constitute malpractice, the broader lesson is stark --- a reward function that measures how legal a response \emph{looks} will produce a model that is maximally photogenic and minimally useful.
Chat is not available.
Successful Page Load