Mind the Gap: A Systematic Review and Benchmark of Spanish-Language Legal Question Answering in Latin America
Abstract
Legal NLP benchmarks have driven progress on question answering, judgment prediction, and legal reasoning, yet a significant blind spot persists: despite Latin America comprising 33 countries with over 650 million inhabitants---most within civil-law traditions---there are \emph{no} dedicated legal QA benchmarks for Spanish-speaking jurisdictions. We present a systematic review of 41 papers (2018--2026) showing that all six Latin American contributions focus exclusively on Brazil and Portuguese. To begin closing this gap, we evaluate five frontier LLMs (GPT-5, Claude Sonnet~4.5, GPT-4o Mini, Llama~4 Maverick, Qwen3-235B) in a zero-shot setting on 458 bar-exam-style multiple-choice questions from Bolivia~(302), Colombia~(115), and Puerto Rico~(41). The best model achieves only 79.7\% overall accuracy despite these questions originating from publicly available materials likely present in pre-training corpora. Performance varies sharply by jurisdiction and legal area, with Criminal Law (66.1\%) and Administrative Law (60.0\%) proving most challenging. These results suggest that current LLMs still face substantial challenges in legal reasoning for civil-law systems and underscore the need for dedicated Spanish-language legal benchmarks beyond multiple-choice formats.