Korean Legal Texts Are Not Legally Korean: Legal Style Refinement and Its Effect on LLM Performance
Abstract
Legal texts across jurisdictions are characterized by linguistic complexity, which poses barriers to comprehension for both human readers and artificial intelligence (AI) systems. In Korea, this challenge is compounded by the historical adoption of law through indirect translation, resulting in a statutory language that systematically deviates from standard modern Korean. We constructed a parallel corpus of 1,105 Korean Civil Code article pairs and fine-tuned A.X-4.0 (72B) to build a legal style-refinement model that rewrites statutory sentences into plain-language equivalents while preserving legal meaning, achieving a mean semantic similarity of 0.96 between model-generated and expert-proposed outputs. By evaluating six state-of-the-art LLMs on Korean legal reasoning benchmarks before and after refinement, we found consistent performance gains of up to 2.03 percentage points for models with higher Korean language and legal knowledge scores, whereas models with lower scores showed slight declines. These findings suggest that the linguistic accessibility of the input text and model-level competence are complementary conditions for reliable legal AI.