Text Enhancement For Paragraph Processing In End-to-end Code-switching TTS
2022 Β· Chunyu Qiang, Jianhua Tao, Ruibo Fu, et al.
Abstract
Current end-to-end code-switching Text-to-Speech (TTS) can already generate high quality two languages speech in the same utterance with single speaker bilingual corpora. When the speakers of the bilingual corpora are different, the naturalness and consistency of the code-switching TTS will be poor. The cross-lingual embedding layers structure we proposed makes similar syllables in different languages relevant, thus improving the naturalness and consistency of generated speech. In the end-to-end code-switching TTS, there exists problem of prosody instability when synthesizing paragraph text. The text enhancement method we proposed makes the input contain prosodic information and sentence-level context information, thus improving the prosody stability of paragraph text. Experimental results demonstrate the effectiveness of the proposed methods in the naturalness, consistency, and prosody stability. In addition to Mandarin and English, we also apply these methods to Shanghaiese and Canto
Authors
(none)
Tags
Stats
Related papers
- Improving Prosody Modelling With Cross-utterance BERT Embeddings For End-to-end Speech Synthesis (2020)10.61
- Towards Natural Bilingual And Code-switched Speech Synthesis Based On Mix Of Monolingual Recordings And Cross-lingual Voice Conversion (2020)0.00
- Paratts: Learning Linguistic And Prosodic Cross-sentence Information In Paragraph-based TTS (2022)8.82
- Contextspeech: Expressive And Efficient Text-to-speech For Paragraph Reading (2023)5.84
- Building A Mixed-lingual Neural TTS System With Only Monolingual Data (2019)0.00
- Diclet-tts: Diffusion Model Based Cross-lingual Emotion Transfer For Text-to-speech -- A Study Between English And Mandarin (2023)9.92
- Cross-lingual Multi-speaker Text-to-speech Synthesis For Voice Cloning Without Using Parallel Corpus For Unseen Speakers (2019)0.00
- Data Processing For Optimizing Naturalness Of Vietnamese Text-to-speech System (2020)4.52