VCTK
Canonical37papers using it
2021first seen
The CSTR VCTK Corpus includes speech data uttered by 110 English speakers with various accents.
Papers using VCTK (37)
- FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with
Single-Step Flow MatchingFSC-Net: Integrating Fast Fourier Convolutions and Progressive Learning for Speech Bandwidth ExtensionCollective Learning Mechanism based Optimal Transport Generative
Adversarial Network for Non-parallel Voice ConversionComplexDec: A Domain-robust High-fidelity Neural Audio Codec with
Complex Spectrum ModelingMambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion ControlEntropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech CodecCausal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2Audio Super-Resolution with Latent Bridge ModelsFNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech SynthesisA High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal LossSpeaker Disentanglement of Speech Pre-trained Model Based on InterpretabilityBridge-SR: Schr\"odinger Bridge for Efficient SRYourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice
Conversion for everyoneStyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion
and Adversarial Training with Large Speech Language ModelsTriAAN-VC: Triple Adaptive Attention Normalization for Any-to-Any Voice
ConversionHiFi-Codec: Group-residual Vector quantization for High Fidelity Audio
CodecFluentSpeech: Stutter-Oriented Automatic Speech Editing with
Context-Aware Diffusion ModelsVALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text
to Speech SynthesizersResGrad: Residual Denoising Diffusion Probabilistic Models for Text to
SpeechSpeech Representation Disentanglement with Adversarial Mutual
Information Learning for One-shot Voice ConversionSelf-Supervised Speech Representations Preserve Speech Characteristics
while Anonymizing VoicesSpeaker Anonymization with Phonetic Intermediate RepresentationsTraining Robust Zero-Shot Voice Conversion Models with Self-supervised
FeaturesCampNet: Context-Aware Mask Prediction for End-to-End Text-Based Speech
EditingRobust Disentangled Variational Speech Representation Learning for
Zero-shot Voice ConversionTowards Improved Zero-shot Voice Conversion with Conditional DSVAEAdapter-Based Extension of Multi-Speaker Text-to-Speech Model for New
SpeakersALO-VC: Any-to-any Low-latency One-shot Voice ConversionFluentEditor: Text-based Speech Editing by Considering Acoustic and
Prosody ConsistencyUnconstrained Dysfluency Modeling for Dysfluent Speech Transcription and
DetectionSingle-channel speech enhancement using learnable loss mixupAttentionStitch: How Attention Solves the Speech Editing ProblemMulti-speaker Text-to-speech Training with Speaker Anonymized DataGLOBE: A High-quality English Corpus with Global Accents for Zero-shot
Speaker Adaptive Text-to-SpeechDreamVoice: Text-Guided Voice ConversionDPSNN: Spiking Neural Network for Low-Latency Streaming Speech
EnhancementFluentEditor2: Text-based Speech Editing by Modeling Multi-Scale
Acoustic and Prosody Consistency