TriBench-Ko: Evaluating LLM Risks in Judicial Workflows
Abstract
Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks focus on proxy tasks such as bar examinations or judgment prediction, offering limited visibility into deployment risks in judicial workflows. To address this, we release TriBench-Ko, a Korean benchmark designed to evaluate potential deployment risks of LLMs within the context of verified judicial task requirements. It covers four core tasks: jurisprudence summarization, precedent retrieval, legal issue extraction, and evidence analysis. It evaluates model behavior across multiple deployment risks including inaccuracy (hallucination, omission, statutory misapplication), biases (demographic, overcompliance), inconsistencies (prompt sensitivity, non-determinism), and adjudicative overreach. Our evaluation of a range of contemporary LLMs reveals that many models exhibit substantial risks particularly in precedent retrieval and omission of legally constitutive information.