Abstract
Expressive voice conversion (VC) aims to transfer the speaker identity and emotion from the target speech to the source speech. In this paper, we propose an expressive voice conversion model with controllable emotional intensity named CEIVC, which includes a specific attribute augmentation (SAA) training strategy for perturbing the speaker and emotional attributes of source speech, and an emotional disentanglement and intensity control (EDIC) module to control the emotional intensity of converted speech. Moreover, we leverage perturbation adaptive instance normalization (PbAdaIN) to further boost the synthesis quality. Experimental results demonstrate that our proposed CEI-VC outperforms five compared models and is capable of controlling the emotional intensity of converted speech. Demos and codes are available at https://tengnn.github.io/ExpressiveVC/.