C-QIRF: An Ordinal Risk Annotation Framework for Contextual Quasi-Identifier Risk in Court Judgments
Abstract
Research on court judgment de-identification has primarily focused on detecting and masking identifiers such as names, addresses, and phone numbers. However, even after identifiers are removed, various pieces of identifiable information may remain when the surrounding context is taken into account. Such information, often referred to as quasi-identifiers, is regulated separately from identifiers in many jurisdictions and therefore requires a distinct layer of de-identification. This paper explores this possibility by proposing C-QIRF (Contextual Quasi-Identifier Risk Framework), a framework for evaluating residual quasi-identifier risk after identifier removal in first-instance divorce judgments from Korean courts using a 0-3 ordinal risk scale. C-QIRF is not a predictive model that estimates the actual probability of re-identification; rather, it is a labeling framework for consistently evaluating candidate expressions and minimal context in pre-release review and subsequent empirical validation. This paper presents the candidate span construction procedure, risk taxonomy, large language model (LLM)-assisted evaluation procedure, and follow-up validation design.