← all papers · overview

QE-XVC: Zero-Shot Cross-Lingual Voice Conversion via Query-Enhancement and Conditional Flow Matching

Abstract

Voice conversion (VC) aims to modify the speaker identity of speech while preserving its linguistic content. Cross-lingual voice conversion (XVC) further enables converting speech between speakers of different languages. Existing zero-shot XVC methods often struggle with pronunciation errors, insufficient speaker modeling or the difficulty of aligning source and target across languages. In this paper, we propose QE-XVC, an efficient zero-shot XVC method with query-enhancement and conditional flow matching (CFM). Specifically, we designed a query-enhancement module in which the speaker embeddings from a speaker verification (SV) model and the content representation are successively used as queries to enhance the frame-level speaker representations extracted by wavelet convolutions. The CFM model then generates converted acoustic features conditioned on both the content and the fine-grained speaker representations, which are subsequently transformed into waveforms by a vocoder. Experimental results demonstrate that, QE-XVC reduces potential pronunciation errors without introducing any quantization or clustering operations, while also improving prosody preservation from the source speaker.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).