Mscenespeech: A Multi-scene Speech Dataset For Expressive Speech Synthesis
2024 Β· Qian Yang, Jialong Zuo, Zhe Su, et al.
Abstract
We introduce an open source high-quality Mandarin TTS dataset MSceneSpeech (Multiple Scene Speech Dataset), which is intended to provide resources for expressive speech synthesis. MSceneSpeech comprises numerous audio recordings and texts performed and recorded according to daily life scenarios. Each scenario includes multiple speakers and a diverse range of prosodic styles, making it suitable for speech synthesis that entails multi-speaker style and prosody modeling. We have established a robust baseline, through the prompting mechanism, that can effectively synthesize speech characterized by both user-specific timbre and scene-specific prosody with arbitrary text input. The open source MSceneSpeech Dataset and audio samples of our baseline are available at https://speechai-demo.github.io/MSceneSpeech/.
Authors
(none)
Tags
Stats
Related papers
- Mntts2: An Open-source Multi-speaker Mongolian Text-to-speech Synthesis Dataset (2022)5.81
- EMOVIE: A Mandarin Emotion Speech Dataset With A Simple Emotional Text-to-speech Model (2021)0.00
- Storytts: A Highly Expressive Text-to-speech Dataset With Rich Textual Expressiveness Annotations (2024)3.58
- Mntts: An Open-source Mongolian Text-to-speech Synthesis Dataset And Accompanied Baseline (2022)5.24
- AISHELL-3: A Multi-speaker Mandarin TTS Corpus And The Baselines (2020)0.00
- Wenetspeech4tts: A 12,800-hour Mandarin TTS Corpus For Large Speech Generation Model Benchmark (2024)9.76
- Speechdialoguefactory: Generating High-quality Speech Dialogue Data To Accelerate Your Speech-llm Development (2025)0.00
- Speechcraft: A Fine-grained Expressive Speech Dataset With Natural Language Description (2024)7.81