A Benchmark for Long-Horizon Reasoning with Dispute Resolution
Abstract
Applying large language models (LLMs) to legal applications has significant social implications. While law is a domain that could significantly benefit from recent advances in LLMs, evaluating these systems remains an open challenge, at the core to which is measuring proficiency in legal reasoning, due to its inherent importance for legal tasks. In this paper, we introduce new datasets for evaluating the quality of execution in open-ended, long-horizon legal research tasks grounded in real-world disputes and their resolutions. We propose ten expert-informed criteria that characterize a high-quality legal answer, covering dimensions such as factual accuracy, legal relevance, reasoning quality, citation support, and practical usefulness. We develop test datasets and synthetically expand them through controlled ablations of individual facts, enabling fine-grained analysis of model robustness and sensitivity to case-specific evidence. Based on this, we establish the performance of frontier LLMs across multiple environments, providing comparative insights into their strengths and weaknesses in the complex legal research setting.