Any-to-one Sequence-to-sequence Voice Conversion Using Self-supervised Discrete Speech Representations
2020 Β· Wen-Chin Huang, Yi-Chiao Wu, Tomoki Hayashi, et al.
Abstract
We present a novel approach to any-to-one (A2O) voice conversion (VC) in a sequence-to-sequence (seq2seq) framework. A2O VC aims to convert any speaker, including those unseen during training, to a fixed target speaker. We utilize vq-wav2vec (VQW2V), a discretized self-supervised speech representation that was learned from massive unlabeled data, which is assumed to be speaker-independent and well corresponds to underlying linguistic contents. Given a training dataset of the target speaker, we extract VQW2V and acoustic features to estimate a seq2seq mapping function from the former to the latter. With the help of a pretraining method and a newly designed postprocessing technique, our model can be generalized to only 5 min of data, even outperforming the same model trained with parallel data.
Authors
(none)
Tags
Stats
Related papers
- S2VC: A Framework For Any-to-any Voice Conversion With Self-supervised Pretrained Representations (2021)12.25
- ACE-VC: Adaptive And Controllable Voice Conversion Using Explicitly Disentangled Self-supervised Speech Representations (2023)0.00
- Non-parallel Sequence-to-sequence Voice Conversion With Disentangled Linguistic And Speaker Representations (2019)14.02
- AAS-VC: On The Generalization Ability Of Automatic Alignment Search Based Non-autoregressive Sequence-to-sequence Voice Conversion (2023)0.00
- Convs2s-vc: Fully Convolutional Sequence-to-sequence Voice Conversion (2018)12.68
- Atts2s-vc: Sequence-to-sequence Voice Conversion With Attention And Context Preservation Mechanisms (2018)14.15
- DRVC: A Framework Of Any-to-any Voice Conversion With Self-supervised Learning (2022)9.59
- Zero-shot Voice Conversion Via Self-supervised Prosody Representation Learning (2021)6.34