Building Reward Models for Grounded Legal Reasoning
Rilton Franzone ⋅ Valentin NOËL ⋅ Puyu Wang ⋅ Phil Torr ⋅ Fabio Fehr
Abstract
Large language models are increasingly used in high-stakes domains such as law, where systems must remain grounded in evidence and abstain when information is insufficient. However, current models frequently hallucinate, over-rely on parametric knowledge, and fail to reliably use retrieved context in retrieval-augmented generation (RAG) settings. We propose a reproducible framework for contextual reward modelling in legal reasoning, constructing preference datasets from legal QA benchmarks capturing answerability, faithfulness, completeness, and legal correctness. We evaluate and train open-source judge models (0.5B–27B) under modest compute constraints. Experiments on legal retrieval and reasoning benchmarks show that DPO fine-tuning on our contextual preference corpus improves overall grounded evaluation up to $+25.5$\,pp compared to general-purpose judges, particularly under insufficient retrieval conditions. Our results demonstrate a step towards reliable legal evaluation without frontier-scale systems, providing a accessible foundation for grounded AI in law and other high-stakes domains.
Chat is not available.
Successful Page Load