LJSpeech
Canonical37papers using it
2021first seen
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books. A transcription is provided for each clip. Clips vary in length from 1 to 10 seconds and have a total length of approximately 24 hours.
Papers using LJSpeech (37)
- Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applicationsAdaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTSClip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality
Text-to-Speech Method based on Contextual Semantic UnderstandingMambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion ControlBeyond Two-stage Diffusion TTS: Joint Structure and Content Refinement via Jump DiffusionAssessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 PerturbationECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency TuningFNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech SynthesisNaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level
QualityStyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion
and Adversarial Training with Large Speech Language ModelsGuided-TTS: A Diffusion Model for Text-to-Speech via Classifier GuidanceDiffVoice: Text-to-Speech with Latent DiffusionUnsupervised word-level prosody tagging for controllable speech
synthesisResGrad: Residual Denoising Diffusion Probabilistic Models for Text to
SpeechEnhancing Gappy Speech Audio Signals with Generative Adversarial
NetworksRepresentative Subset Selection for Efficient Fine-Tuning in
Self-Supervised Speech RecognitionLow-Resource Text-to-Speech Using Specific Data and Noise AugmentationEnergy-Based Models For Speech SynthesisSchrodinger Bridges Beat Diffusion Models on Text-to-Speech SynthesisFederated Learning with Dynamic Transformer for Text to SpeechJETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to
SpeechU-DiT TTS: U-Diffusion Vision Transformer for Text-to-SpeechRep2wav: Noise Robust text-to-speech Using self-supervised
representationsHiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise
Filter and Inverse Short Time Fourier TransformIndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech
Corpus for Scaling Indian TTSContinual Learning in Machine Speech Chain Using Gradient Episodic
MemoryMixer-TTS: non-autoregressive, fast and compact text-to-speech model
conditioned on language model embeddingsSOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural
Text-to-Speech SynthesisCross-Utterance Conditioned VAE for Non-Autoregressive Text-to-SpeechImproved Consistency Training for Semi-Supervised Sequence-to-Sequence
ASR via Speech Chain Reconstruction and Self-TranscribingUnsupervised ASR via Cross-Lingual Pseudo-LabelingArabic Dysarthric Speech Recognition Using Adversarial and Signal-Based
AugmentationReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-SpeechAttentionStitch: How Attention Solves the Speech Editing ProblemLlama-VITS: Enhancing TTS Synthesis with Semantic AwarenessSequence-to-sequence models in peer-to-peer learning: A practical
applicationPRESENT: Zero-Shot Text-to-Prosody Control