Converting Anyone's Voice: End-to-end Expressive Voice Conversion With A Conditional Diffusion Model
2024 Β· Zongyang Du, Junchen Lu, Kun Zhou, et al.
Abstract
Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC has not been extensively explored. Previous approaches have relied on vocoders for speech reconstruction, which makes speech quality heavily dependent on the performance of vocoders. A major challenge of expressive VC lies in emotion prosody modeling. To address these challenges, this paper proposes a fully end-to-end expressive VC framework based on a conditional denoising diffusion probabilistic model (DDPM). We utilize speech units derived from self-supervised speech models as content conditioning, along with deep features extracted from speech emotion recognition and speaker verification systems to model emotional style and speaker identity. Objective and subjective evaluations show the effectiveness of our framework. Codes and samples are publicly available.
Authors
(none)
Tags
Stats
Related papers
- Expressive Voice Conversion: A Joint Framework For Speaker Identity And Emotional Style Transfer (2021)9.03
- Disentanglement Of Emotional Style And Speaker Identity For Expressive Voice Conversion (2021)10.97
- Emoreg: Directional Latent Vector Modeling For Emotional Intensity Regularization In Diffusion-based Voice Conversion (2024)2.26
- PMVC: Data Augmentation-based Prosody Modeling For Expressive Voice Conversion (2023)9.23
- ZSDEVC: Zero-shot Diffusion-based Emotional Voice Conversion With Disentangled Mechanism (2024)0.00
- Expressive-vc: Highly Expressive Voice Conversion With Attention Fusion Of Bottleneck And Perturbation Features (2022)9.03
- DDDM-VC: Decoupled Denoising Diffusion Models With Disentangled Representation And Prior Mixup For Verified Robust Voice Conversion (2023)11.29
- Codiff-vc: A Codec-assisted Diffusion Model For Zero-shot Voice Conversion (2024)0.00