LibriTTS
Emerging44papers using it
2021first seen
LibriTTS is a dataset that contains high-quality, diverse speech recordings used to evaluate text-to-speech systems' performance in generating natural-sounding speech.
Papers using LibriTTS (44)
- Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned
Flow Matching ModelLibriConvo: Simulating Conversations from Read Literature for ASR and DiarizationVoCodec: A Low-bitrate Streamable Neural Speech Codec with Voicing-driven QuantizationMambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion ControlEntropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech CodecZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech SynthesisCausal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2Principled Coarse-Grained Acceptance for Speculative Decoding in SpeechLlasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech SynthesisFNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech SynthesisRephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style TransferPseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech SynthesisTowards Flow-Matching-based TTS without Classifier-Free GuidanceSSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech
Editing and SynthesisBigVGAN: A Universal Neural Vocoder with Large-Scale TrainingStyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion
and Adversarial Training with Large Speech Language ModelsHiFi-Codec: Group-residual Vector quantization for High Fidelity Audio
CodecDiffVoice: Text-to-Speech with Latent DiffusionFluentSpeech: Stutter-Oriented Automatic Speech Editing with
Context-Aware Diffusion ModelsCML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource
LanguagesResGrad: Residual Denoising Diffusion Probabilistic Models for Text to
SpeechOn granularity of prosodic representations in expressive text-to-speechInterpretable Style Transfer for Text-to-Speech with ControlVAE and
Diffusion BridgeRep2wav: Noise Robust text-to-speech Using self-supervised
representationsDistinguishing Neural Speech Synthesis Models Through Fingerprints in
Speech WaveformsHiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise
Filter and Inverse Short Time Fourier TransformEnhancing the Stability of LLM-based Speech Generation Systems through
Self-Supervised RepresentationsIndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech
Corpus for Scaling Indian TTSTraining Robust Zero-Shot Voice Conversion Models with Self-supervised
FeaturesCampNet: Context-Aware Mask Prediction for End-to-End Text-Based Speech
EditingCross-Utterance Conditioned VAE for Non-Autoregressive Text-to-SpeechGlow-WaveGAN 2: High-quality Zero-shot Text-to-speech Synthesis and
Any-to-any Voice ConversionCan we use Common Voice to train a Multi-Speaker TTS system?Adapter-Based Extension of Multi-Speaker Text-to-Speech Model for New
SpeakersLibriTTS-R: A Restored Multi-Speaker Text-to-Speech CorpusCross-Utterance Conditioned VAE for Speech GenerationGLOBE: A High-quality English Corpus with Global Accents for Zero-shot
Speaker Adaptive Text-to-SpeechDreamVoice: Text-Guided Voice ConversionOn the Effectiveness of Acoustic BPE in Decoder-Only TTSAccelerating High-Fidelity Waveform Generation via Adversarial Flow
Matching OptimizationDisentangling the Prosody and Semantic Information with Pre-trained
Model for In-Context Learning based Zero-Shot Voice ConversionAnalyzing and Mitigating Inconsistency in Discrete Audio Tokens for
Neural Codec Language ModelsFluentEditor2: Text-based Speech Editing by Modeling Multi-Scale
Acoustic and Prosody ConsistencyLibri2Vox Dataset: Target Speaker Extraction with Diverse Speaker
Conditions and Synthetic Data