The X-LANCE Technical Report For Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge
2024 Β· Yiwei Guo, Chenrun Wang, Yifan Yang, et al.
Abstract
Discrete speech tokens have been more and more popular in multiple speech processing fields, including automatic speech recognition (ASR), text-to-speech (TTS) and singing voice synthesis (SVS). In this paper, we describe the systems developed by the SJTU X-LANCE group for the TTS (acoustic + vocoder), SVS, and ASR tracks in the Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge. Notably, we achieved 1st rank on the leaderboard in the TTS track both with the whole training set and only 1h training data, with the highest UTMOS score and lowest bitrate among all submissions.
Authors
(none)
Tags
Stats
Related papers
- UTDUSS: Utokyo-sarulab System For Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge (2024)0.00
- The SJTU X-LANCE Lab System For CNSRC 2022 (2022)6.22
- Transformer VQ-VAE For Unsupervised Unit Discovery And Speech Synthesis: Zerospeech 2020 Challenge (2020)9.41
- The Zero Resource Speech Challenge 2019: TTS Without T (2019)13.17
- The Zero Resource Speech Challenge 2020: Discovering Discrete Subword And Word Units (2020)11.58
- Evaluating Text-to-speech Synthesis From A Large Discrete Token-based Speech Language Model (2024)0.00
- Mulantts: The Microsoft Speech Synthesis System For Blizzard Challenge 2023 (2023)5.84
- Transsion Tsup's Speech Recognition System For ASRU 2023 MADASR Challenge (2023)0.00