Autoregressive Zero-Shot Voice Conversion
Luca Lanzendörfer ⋅ Frédéric Berdoz ⋅ Antonis Asonitis ⋅ Roger Wattenhofer
Abstract
We present LaVoco, an autoregressive model for zero-shot voice conversion built on a pretrained autoregressive text-to-speech backbone. LaVoco leverages features extracted from Whisper and XCodec2 discrete tokens to autoregressively generate the target utterance using a reference timbre. We analyze different model variants of LaVoco and find that our dual representation achieves the best content preservation and competitive speaker similarity while retaining the scalability of next-token prediction. Our approach also proves robust to sub-1 second reference timbre prompts, maintaining high speaker similarity. Finally, we release our code, pretrained models, and a 10k-hour dataset of synthetic speech pairs to support future research.
Video
Chat is not available.
Successful Page Load