Multilingual Speech Editing
Antonis Asonitis ⋅ Luca Lanzendörfer ⋅ Frédéric Berdoz ⋅ Roger Wattenhofer
Abstract
We present MuSE, an autoregressive model for multilingual speech editing at both phoneme and word-level granularity. We train MuSE using a three-stage curriculum learning approach, progressing from text-phoneme grounding to full bidirectional audio infilling. To support training at this scale, we release multilingual-audio-alignments, the largest publicly available word-and-phoneme-level aligned speech corpus to date, spanning over 32k hours across 13 languages. We demonstrate state-of-the-art performance on the established RealEdit benchmark and introduce MultiLingualEdit, a multilingual extension with handcrafted editing examples in 13 languages. The codebase and datasets are made publicly available.
Video
Chat is not available.
Successful Page Load