Scaling Behavior in Model Fine-tuning for Audio DeepFake Detection
Abstract
Recent advances in audio deepfake detection have been driven by increasingly large speech foundation models and growing amounts of synthetic data. Despite strong benchmark performance, it remains unclear how detection capability scales with model capacity and training data under realistic deployment conditions involving distribution shift, signal corruption, and unseen synthesis pipelines. In this work, we present the first systematic study of scaling laws in post-training audio deepfake detection, focusing on fine-tuning regimes rather than large-scale pretraining. Using a controlled family of speech foundation models with shared architecture and pretraining, we analyze how detection performance, robustness, and generalization evolve as a function of model size and training data scale. Our results reveal a fundamental asymmetry between performance scaling and robustness scaling in audio deepfake detection, suggesting increasing model capacity alone is insufficient for achieving reliable real-world generalization.
Lay Summary
AI can now create highly realistic fake voices, raising concerns about scams, misinformation, and identity impersonation. To address this problem, researchers have developed audio deepfake detectors that try to tell whether speech is real or AI-generated. In this work, we study how these detectors improve as models and training datasets become larger. We test them under challenging real-world conditions, including noisy audio, different languages, and unseen voice generation systems. We find that larger models generally perform better, but their improvements are much smaller in difficult real-world settings. Even the largest models still struggle with unfamiliar or corrupted audio. Our results show that simply scaling up model size and data is not enough to build reliable audio deepfake detectors for real-world use.