Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases
Abstract
Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain underexplored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We benchmark OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what accounts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, while the human evaluation is overall stricter compared to the ones from models (LLM-as-Judge). Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not lead to more accurate predictions, compared to the rest of the examined settings. Based on our findings, we urge the community not to rely solely on automated evaluation and to avoid the use of task accuracy as an appropriate proxy of reasoning quality.