Crossvoice: Crosslingual Prosody Preserving Cascade-s2st Using Transfer Learning
2024 Β· Medha Hira, Arnav Goel, Anubha Gupta
Abstract
This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We conducted comprehensive experiments comparing CrossVoice with direct-S2ST systems, showing improved BLEU scores on tasks such as Fisher Es-En, VoxPopuli Fr-En and prosody preservation on benchmark datasets CVSS-T and IndicTTS. With an average mean opinion score of 3.75 out of 4, speech synthesized by CrossVoice closely rivals human speech on the benchmark, highlighting the efficacy of cascade-based systems and transfer learning in multilingual S2ST with prosody transfer.
Authors
(none)
Tags
Stats
Related papers
- Transvip: Speech To Speech Translation System With Voice And Isochrony Preservation (2024)5.24
- Cross-lingual Text-to-speech With Flow-based Voice Conversion For Improved Pronunciation (2022)0.00
- Cross-lingual Knowledge Distillation Via Flow-based Voice Conversion For Robust Polyglot Text-to-speech (2023)0.00
- Building Multi Lingual TTS Using Cross Lingual Voice Conversion (2020)0.00
- Voice Conversion By Cascading Automatic Speech Recognition And Text-to-speech Synthesis With Prosody Transfer (2020)5.84
- Crossspeech: Speaker-independent Acoustic Representation For Cross-lingual Speech Synthesis (2023)7.16
- Translatotron 2: High-quality Direct Speech-to-speech Translation With Voice Preservation (2021)0.00
- Towards Natural And Controllable Cross-lingual Voice Conversion Based On Neural TTS Model And Phonetic Posteriorgram (2021)0.00