Split Group Knockoffs: Controlling False Discovery Rate in Transformational Group Sparsity
Abstract
Controlling the false discovery rate (FDR) under complex sparsity structures remains a fundamental challenge in large language model (LLM) analysis. Motivated by multiple comparison problems in LLMs, we consider a setting in which sparsity arises at the group level after a linear transformation of model parameters. We propose Split Group Knockoffs (SGKs), a general framework for group-wise variable selection under grouped transformational sparsity that extends the Split Knockoff procedure to grouped transformed variables. We establish theoretical guarantees for group-level FDR control and support recovery consistency, addressing challenges induced by group-wise penalties in transformed spaces. Applying SGK to LLM behavior auditing experiment reveals that model disagreement is not uniform across subjects, but instead concentrates in domains with greater semantic and reasoning complexity, where SGK effectively distinguishes genuine behavioral deviations from surface-level performance variation.
Lay Summary
Large language models are often compared using overall benchmark scores, but these averages can hide important differences in how models behave across subjects such as physics, medicine, or psychology. Existing statistical methods are not designed to reliably identify these differences when model behaviors are organized into complex groups and many comparisons are made simultaneously. To address this problem, we develop Split Group Knockoffs (SGK), a statistical framework that detects meaningful group-level behavioral deviations while controlling false discoveries. We apply SGK to auditing state-of-the-art language models on the MMLU-Pro benchmark and show that model disagreements are concentrated in subjects requiring more complex reasoning, rather than appearing uniformly across all domains. We also apply the method to Alzheimer’s disease brain imaging data, where it identifies brain regions and neural connections associated with disease progression. Our results demonstrate that SGK can help researchers distinguish genuine patterns from random variation in both AI systems and scientific data analysis.