The Legal OCR Trust Gap: A Dual-Model Audit of OCR-Augmented Vision-Language Models on Contracts, Forms, and Korean Criminal Cases, Mechanistically Traced to Four Attention Heads
Karim Mahfouz
Abstract
In Mata v. Avianca (S.D.N.Y. 2023), attorneys were sanctioned for submitting AI-generated fictional case citations they had not verified; the resulting professional-responsibility doctrine now requires lawyers to understand the limitations of the AI tools they deploy. The standard pipeline for AI contract review feeds a vision-language model (VLM) both a document image and a separately produced OCR transcript—a deployment pattern whose mechanism-level failure modes have not been characterized. Legal practice depends on this pipeline being faithful to the original document. We audit it across three real legal benchmarks and document a conditional failure: OCR injection produces fluent, confident, professional-sounding output that does not depend on the image. On CUAD commercial contracts ($N=217$ paired items, 1,302 predictions per model), clean PaddleOCR injection produces a Bonferroni-significant trust gap on both Qwen2.5-VL-7B (+10.60pp substring, $p=6.2\times 10^{-4}$, McNemar $b=34/c=11$, OR=3.09) and InternVL3-8B (+30.41pp substring, +15.35pp ANLS, McNemar $b=71/c=5$, OR=14.2, $p<10^{-13}$). Critically, three-character adversarial edits to numerical clauses produce statistically indistinguishable gains on both models (+10.14pp and +29.49pp; McNemar nulls $p=1.00$ and $p=0.73$): adversarial OCR is treated as authoritative. The aggregate trust gap is a net average of opposing effects: OCR injection hurts items the model could already solve (Qwen $-8.53$pp on $n=129$; InternVL3 $-7.58$pp) while dramatically rescuing items it could not (Qwen +38.64pp on $n=88$; InternVL3 +47.02pp on $n=151$). On FUNSD scanned forms (pooled $N=180$), clean OCR decreases accuracy by 5.22pp ($p=3.5\times 10^{-7}$); scrambled OCR is ignored. On Korean LBox criminal-case excerpts (pooled $N=180$), Korean-OCR injection collapses the Hangul fraction of Qwen's output from 0.84 to 0.33 ($p=2.3\times 10^{-22}$): the model produces English summaries of Korean defendants and Korean courts. We localize the mechanism to four attention heads in layers 0–1 of Qwen's language backbone (L0H9, L0H17, L0H20, L1H2). Ablating these four heads rescues Korean output by +30.4pp Hangul on paired $N=30$ (Cohen's $d=+0.650$, medium); the full mechanism intervention (no OCR + ablation) recovers +63.3pp substring with McNemar $b=19/c=0$ ($p=3.8\times 10^{-6}$, Cohen's $d=+1.292$, huge). A size-matched random-head control moves in the opposite direction. The intervention does not transfer to InternVL3-8B, indicating the mechanism is wrapper-specific to the Qwen2.5-VL family. We map all three failures to EU AI Act Article 15 conformity obligations, ABA Formal Opinion 512 competence duties, and NIST AI RMF measurement controls, and provide three concrete checks for procurement officers. All raw predictions, statistics, and code are released under permissive license.
Chat is not available.
Successful Page Load