← all papers · overview

Marco-Voice: A Unified Framework for Expressive Speech Synthesis with Voice Cloning

Abstract

This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly expressive, controllable, and natural speech generation that faithfully preserves speaker identity across diverse linguistic and emotional contexts. To enable independent manipulation of speaker identity and emotion, we propose an effective orthogonal speaker-emotion disentanglement and in-batch contrastive learning method, along with a rotational emotional embedding integration method for smooth emotion control. We also introduce a cross-attention module between emotional representations and speech tokens to better integrate emotional information with acousCosyVoice1 [3], a scalable multitic content. To support comprehensive training and evaluation, we construct CSEMOTIONS, a Chinese high-quality emotional speech dataset containing 10 hours of Mandarin speech from ten professional speakers across seven emotional categories. Extensive experiments show that Marco-Voice achieves significant improvements in both objective and subjective metrics, demonstrating superior performance in speaker similarity and emotional richness and marking a substantial advance in expressive neural speech synthesis. Our code and dataset are publicly available at https://github.com/AIDC-AI/Marco-Voice and https://huggingface.co/datasets/AIDC-AI/CSEMOTIONS respectively.

Code

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).