Estimating Model-Level Membership Inference Vulnerability Without Reference Models
Euodia Dodd ⋅ Nataša Krčo ⋅ Igor Shilov ⋅ Matthew Wicker ⋅ Yves-Alexandre de Montjoye
Abstract
Membership inference attacks (MIAs) are the standard tool for evaluating the privacy risks of AI models, but state-of-the-art attacks require training tens to hundreds of expensive reference models. We present a framework for estimating model-level vulnerability to the Likelihood Ratio Attack (LiRA) directly from the train and test loss distributions of the target model, with no reference models required. We show that LiRA's per-sample signal decomposes into a variance-ratio term and a residual mean-shift term, placing models on a continuum of uncertainty collapse that is directly observable from loss distribution shape. At the heavy-tailed end (image classifiers), the LOSS attack TNR predicts LiRA TPR@FPR$=10^{-3}$ with RMSE 0.03 across 9 architectures and 4 datasets, outperforming low-cost reference-model attacks such as RMIA. At the symmetric end (LLMs), the LOSS attack AUC predicts LiRA TPR with RMSE 0.01 across five GPT-2 sizes from 10M to 1B parameters.
Chat is not available.
Successful Page Load