AISHELL-3
Emerging7papers using it
2022first seen
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It can be used to train multi-speaker Text-to-Speech (TTS) systems.The corpus contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers and
Papers using AISHELL-3 (7)
- M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech SynthesisEnhanced exemplar autoencoder with cycle consistency loss in any-to-one
voice conversionLanguage-Independent Speaker Anonymization Approach using
Self-Supervised Pre-Trained ModelsAdaVocoder: Adaptive Vocoder for Custom VoicePSVRF: Learning to restore Pitch-Shifted Voice without referenceTGAVC: Improving Autoencoder Voice Conversion with Text-Guided and
Adversarial TrainingPMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice
Conversion