Jurisdiction-Sensitive Legal AI: Evaluating Legal Competence as Structure Preservation
Fryderyk Kuzma
Abstract
Legal AI is commonly assessed as answer accuracy: a system is judged correct when its conclusion matches a gold label. We argue that this target is structurally inadequate, because the legal acceptance, reviewability, applicability, or procedural availability of a conclusion depends on context, sources, authority ordering, defeaters, burdens of proof, procedure, temporal validity, and language. We define $\COJ$, a jurisdiction-indexed structured argument record emitted by a system, and $\VLOJ$, the subset of candidate outputs whose fields satisfy jurisdiction-specific validation predicates. We model answer-only scoring as a structure-forgetting projection from $\COJ$ to the bare conclusion $q$; this projection is many-to-one, has no guarantee of lifting answer-level support to a validation-satisfying legal argument, and does not determine whether the underlying argument is accepted, defeated, unresolved, outdated, or procedurally unavailable. We derive a fourteen-criterion evaluation rubric over the candidate fields, introduce jurisdictional structure profiles that describe how different legal systems impose different recovery burdens on AI, and apply the framework qualitatively to documented legal-AI failures and governance actions. The central message is that generalization across legal systems is not porting a conclusion by textual similarity but reconstructing, in each jurisdiction's own terms, a candidate object that satisfies that jurisdiction's validation predicates — or reporting that none exists.
Chat is not available.
Successful Page Load