Learning Explicit Prosody Models And Deep Speaker Embeddings For Atypical Voice Conversion
2020 Β· Disong Wang, Songxiang Liu, Lifa Sun, et al.
Abstract
Though significant progress has been made for the voice conversion (VC) of typical speech, VC for atypical speech, e.g., dysarthric and second-language (L2) speech, remains a challenge, since it involves correcting for atypical prosody while maintaining speaker identity. To address this issue, we propose a VC system with explicit prosodic modelling and deep speaker embedding (DSE) learning. First, a speech-encoder strives to extract robust phoneme embeddings from atypical speech. Second, a prosody corrector takes in phoneme embeddings to infer typical phoneme duration and pitch values. Third, a conversion model takes phoneme embeddings and typical prosody features as inputs to generate the converted speech, conditioned on the target DSE that is learned via speaker encoder or speaker adaptation. Extensive experiments demonstrate that speaker adaptation can achieve higher speaker similarity, and the speaker encoder based conversion model can greatly reduce dysarthric and non-native pronu
Authors
(none)
Tags
Stats
Related papers
- Zero-shot Voice Conversion Via Self-supervised Prosody Representation Learning (2021)6.34
- PMVC: Data Augmentation-based Prosody Modeling For Expressive Voice Conversion (2023)9.23
- ACE-VC: Adaptive And Controllable Voice Conversion Using Explicitly Disentangled Self-supervised Speech Representations (2023)0.00
- Duta-vc: A Duration-aware Typical-to-atypical Voice Conversion Approach With Diffusion Probabilistic Model (2023)0.00
- Converting Anyone's Voice: End-to-end Expressive Voice Conversion With A Conditional Diffusion Model (2024)5.24
- Assem-vc: Realistic Voice Conversion By Assembling Modern Speech Synthesis Techniques (2021)11.64
- Voice Reenactment With F0 And Timing Constraints And Adversarial Learning Of Conversions (2021)2.26
- Speaker Identity Preservation In Dysarthric Speech Reconstruction By Adversarial Speaker Adaptation (2022)0.00