← all papers · overview

Source Speech Reconstruction for Many-to-Many and One-to-One Voice Conversion

Abstract

: As voice conversion (VC) systems grow in realism and accessibility, concerns over their misuse – particularly in deepfake scams – have intensified. In this paper, we demonstrate an approach to reconstructing attacker speech, which is applicable to one-to-one (O2O) and many-to-many (M2M) voice conversion. Unlike previous solutions based on extracting high-dimensional embeddings, we demonstrate that the attacker voice is deeply embedded into the VC model and thus it cannot be removed. Furthermore, we demonstrate that we can subsequently reconstruct their voice using the original VC model in order to link the attacker identity back to the deepfake audio. We show that current O2O VC models enable reconstruction without additional training. For M2M models, we introduce a source speaker classifier to facilitate reconstruction. Empirical evaluation across state-of-the-art models (MaskCycleGAN-VC, StarGANv2-VC) demonstrates that reconstructed speech achieves as low as EER 0.14 % for M2M VC reconstruction and 5.15 % EER for O2O VC reconstruction. As VC models improve, this method offers a scalable path for source speaker attribution in forensic and security applications.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).