Flow Fake: Parametric Efficient Alternative for Transformers
Abstract
Audio deepfakes generated by neural text- to-speech and voice-cloning systems threaten speaker verification and public discourse at scale. The core challenge is cross-dataset generalisa- tion: detectors trained on one synthesis pipeline collapse on unseen forgeries. We argue that this failure is primarily because of structural synthetic- speech artefacts which are multi-timescale trajec- tory anomalies. Though every existing detector aggregates a fixed-window frame statistics, this misaligns the architecture with the signal. We propose LIQNN, a Liquid Time-Constant (LTC) architecture whose hidden state evolves via a learned ODE, with per-neuron adaptive time con- stants simultaneously resolving spectral (∼10 ms) and prosodic (∼2 s) cues. At only ≈34 K parame- ters LIQNN achieves formal BIBO stability and O(∆t4) integration error. On a four-dataset cross- domain benchmark (ASVspoof 2019-LA, Fake- OrReal, InTheWild, MLAAD), LIQNN reaches 75.29% on ASVspoof 2019 trained only on Fake- OrReal and 79.97% trained only on MLAAD. It outperforms RawGAT-ST and Whisper-DF on ev- ery evaluated pair and matching SSL Wav2vec2 (300× larger) at 0.01% of its parameter count.