LegalHalluLens: Typed Hallucination Failure Modes and Calibrated Multi-Agent Debate Mitigation
Abstract
AI systems deployed for legal workflows hallucinate at rates that aggregate metrics report at ∼52%, but this single number conceals which kinds of claim fail and in which direction the resulting bias runs. We present LegalHalluLens, a typed hallucination diagnostic for LLM contract extraction, paired with a six-role multi-agent debate pipeline that uses the diagnosis to calibrate its mitigation. The diagnostic has two parts: typed hallucination profiles across four claim categories (numeric, temporal, obligation, factual), and a Risk Direction Index (RDI) that decomposes content errors into invention versus omission. Across CUAD (Hendrycks et al., 2021) (510 contracts; 249,252 clause-level instances; four architecturally diverse models) we measure a within-model gap of∼38–40 pp between obligation/numeric and temporal claims that aggregate reporting hides, and show that two extractors with matched 52% rates can carry opposite RDIs. The calibrated debate pipeline reduces fabricated detections by 45% with per-category gains tracking the diagnosis, allowing a 4B-active open backbone to rank first under 4 of 5 weighting schemes against commercial frontiers.