Almost Unsupervised Text To Speech And Automatic Speech Recognition
2019 Β· Yi Ren, Xu Tan, Tao Qin, et al.
Abstract
Text to speech (TTS) and automatic speech recognition (ASR) are two dual tasks in speech processing and both achieve impressive performance thanks to the recent advance in deep learning and large amount of aligned speech and text data. However, the lack of aligned data poses a major practical problem for TTS and ASR on low-resource languages. In this paper, by leveraging the dual nature of the two tasks, we propose an almost unsupervised learning method that only leverages few hundreds of paired data and extra unpaired data for TTS and ASR. Our method consists of the following components: (1) a denoising auto-encoder, which reconstructs speech and text sequences respectively to develop the capability of language modeling both in speech and text domain; (2) dual transformation, where the TTS model transforms the text \(y\) into speech \(\hat\{x\}\), and the ASR model leverages the transformed pair \((\hat\{x\},y)\) for training, and vice versa, to boost the accuracy of the two tasks; (3
Authors
(none)
Tags
Stats
Related papers
- Towards Unsupervised Automatic Speech Recognition Trained By Unaligned Speech And Text Only (2018)0.00
- Semi-supervised Sequence-to-sequence ASR Using Unpaired Speech And Text (2019)0.00
- Unsupervised Text-to-speech Synthesis By Unsupervised Automatic Speech Recognition (2022)12.92
- You Do Not Need More Data: Improving End-to-end Speech Recognition By Text-to-speech Data Augmentation (2020)11.49
- Improving Robustness Of Neural Inverse Text Normalization Via Data-augmentation, Semi-supervised Learning, And Post-aligning Method (2023)0.00
- Improving Accented Speech Recognition Using Data Augmentation Based On Unsupervised Text-to-speech Synthesis (2024)0.00
- Unsupervised Learning For Sequence-to-sequence Text-to-speech For Low-resource Languages (2020)9.59
- A General Multi-task Learning Framework To Leverage Text Data For Speech To Text Tasks (2020)11.67