Unsupervised Learning For Sequence-to-sequence Text-to-speech For Low-resource Languages
2020 Β· Haitong Zhang, Yue Lin
Abstract
Recently, sequence-to-sequence models with attention have been successfully applied in Text-to-speech (TTS). These models can generate near-human speech with a large accurately-transcribed speech corpus. However, preparing such a large data-set is both expensive and laborious. To alleviate the problem of heavy data demand, we propose a novel unsupervised pre-training mechanism in this paper. Specifically, we first use Vector-quantization Variational-Autoencoder (VQ-VAE) to ex-tract the unsupervised linguistic units from large-scale, publicly found, and untranscribed speech. We then pre-train the sequence-to-sequence TTS model by using the<unsupervised linguistic units, audio>pairs. Finally, we fine-tune the model with a small amount of<text, audio>paired data from the target speaker. As a result, both objective and subjective evaluations show that our proposed method can synthesize more intelligible and natural speech with the same amount of paired training data. Besides, we extend our
Authors
(none)
Tags
Stats
Related papers
- DQR-TTS: Semi-supervised Text-to-speech Synthesis With Dynamic Quantized Representation (2023)2.26
- QS-TTS: Towards Semi-supervised Text-to-speech Synthesis Via Vector-quantized Self-supervised Speech Representation Learning (2023)2.26
- Learning To Speak From Text: Zero-shot Multilingual Text-to-speech With Unsupervised Text Pretraining (2023)8.82
- Generating Diverse And Natural Text-to-speech Samples Using A Quantized Fine-grained VAE And Auto-regressive Prosody Prior (2020)12.54
- Transfer Learning Framework For Low-resource Text-to-speech Using A Large-scale Unlabeled Speech Corpus (2022)10.21
- Almost Unsupervised Text To Speech And Automatic Speech Recognition (2019)0.00
- Unsupervised Text-to-speech Synthesis By Unsupervised Automatic Speech Recognition (2022)12.92
- VQTTS: High-fidelity Text-to-speech Synthesis With Self-supervised VQ Acoustic Feature (2022)11.85