Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice
Abstract
Legal AI benchmark research will frequently invoke the assumption that large language models can improve access to justice, including for people who cannot access lawyers in order to understand and exercise their legal rights. However, we argue that current benchmarks are not equipped to support this assumption because they evaluate legal reasoning over inputs that have already been preprocessed by legal experts, which measures the upper bound of model performance. Access to justice depends on a lower bound: how models perform when inputs come from pro se litigants, whose prompts may contain noisy narratives, buried facts, omissions, folk-legal assumptions, and surface-level errors. These degradations are comparable to the conditions under which LLMs have been shown to degrade in the general ML literature, such as long context sensitivity, underspecification, hallucination, and typographical perturbations. We connect evidence from pro se literature with this body of ML research and present a small perturbation experiment on LEXam, a leading legal benchmark, to illustrate the gap between the two bounds. If models continue to focus on improvements based on current benchmarks that only measure the upper bound, this gap will continue to remain hidden or even widen. We conclude by calling for legal benchmarks that directly measure robustness under pro se-like inputs so that access to justice claims about legal AI can become empirically testable.