LibriSpeech
Canonical355papers using it
2021first seen
~1,000 hours of read English audiobook speech with transcripts โ a standard ASR benchmark.
Papers using LibriSpeech (200)
- Revisiting ASR Error Correction with Specialized ModelsLCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription RobustnessVoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and ExtrapolationRIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible EvaluationAudio-Conditioned Diffusion LLMs for ASR and Deliberation ProcessingTest-Time Compute Scaling for ASR with Depth-Conditioned Looped TransformersTowards Data-free and Training-free Compression for Speech Foundation Models Using Parameter ClusteringDASH: Dual-View Self-Distillation with Multi-Layer Hidden Representations for Robust Speech RecognitionNeural Speaker Diarization via Multilingual Training: Evaluation on Low-Resource Nepali-Hindi SpeechImproving Streaming Speech Recognition With Time-Shifted Contextual
Attention And Dynamic Right Context MaskingTASLA: Text-Aligned Speech Tokens with Multiple Layer-AggregationStreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive ModelingSegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer
Based Speech RecognitionA Domain Adaptation Framework for Speech Recognition Systems with Only
Synthetic dataA Neural Model for Contextual Biasing Score Learning and FilteringPARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity DisambiguationPAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech RecognitionAnalysis of Domain Shift across ASR Architectures via TTS-Enabled Separation of Target Domain and Acoustic ConditionsZero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-PronunciationHybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence ModelsUnifying Streaming and Non-streaming Zipformer-based ASRBoosting CTC-Based ASR Using LLM-Based Intermediate Loss RegularizationCalm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads DownAttractor-Based Speech Separation of Multiple Utterances by Unknown Number of SpeakersEffective and Efficient Mixed Precision Quantization of Speech
Foundation ModelsWhisfusion: Parallel ASR Decoding via a Diffusion TransformerSpiking and Event-driven Neuromorphic Mamba Models for Efficient Speech RecognitionIs Text All You Need? Text as a Universal Information Bottleneck for Speech LLMsSpeech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and TranslationNon-Autoregressive Minimum Bayes' Risk Decoding for Fast Speech RecognitionBeyond Noise Suppression: Dynamic Distortion Control Loss for Speech Enhancement and Robust Automatic Speech RecognitionA Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASRPhonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASRText-Utilization for Encoder-dominated Speech Recognition ModelsWhisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA ConversionWhisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room AcousticsPhonemeDF: A Synthetic Speech Dataset for Audio Deepfake Detection and Naturalness EvaluationLLMs and Speech: Integration vs. CombinationLearnable Pulse Accumulation for On-Device Speech Recognition: How Much Attention Do You Need?Modeling Overlapped Speech with ShufflesDiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion TransformersFrom Oracle to Noisy Context: Mitigating Contextual Exposure Bias in Speech-LLMsDecoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASROnline Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future ContextIKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech RecognitionASR for Affective Speech: Investigating Impact of Emotion and Speech Generative StrategyPeeking Into The Future For Contextual BiasingSupplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious DatasetSpeech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech RecognitionReal-Time Speech Enhancement via a Hybrid ViT: A Dual-Input Acoustic-Image Feature FusionIDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector MappingSpiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early ExitingTowards Unsupervised Speech Recognition at the Syllable-LevelGDiffuSE: Diffusion-based speech enhancement with noise model guidanceTokenChain: A Discrete Speech Chain via Semantic Token ModelingArticulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention FusionFLToP CTC: Frame-Level Token Pruning via Relative Threshold for Efficient and Memory-Saving Decoding on Diverse PlatformsRLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech SynthesisA Study of the Removability of Speaker-Adversarial PerturbationsNoisy Disentanglement with Tri-stage Training for Noise-Robust Speech RecognitionSpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource SettingsFuseCodec: Semantic-Contextual Fusion and Supervision for Neural CodecsBiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech RecognitionPitch Accent Detection improves Pretrained Automatic Speech RecognitionObjective Soups: Multilingual Multi-Task Modeling for Speech ProcessingJSQA: Speech Quality Assessment with Perceptually-Inspired Contrastive Pretraining Based on JND Audio PairsWhale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech dataLow-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-TuningTraining-Free Voice Conversion with Factorized Optimal TransportAdapting Whisper for Streaming Speech Recognition via Two-Pass DecodingCMU's IWSLT 2025 Simultaneous Speech Translation SystemEnhanced Hybrid Transducer and Attention Encoder Decoder with Text DataWTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues PeservationTowards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data AugmentationTowards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic PrecisionEvaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech RecognitionContextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and ActivationAn Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASRA Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal TransportMTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition SystemsExploring Gender Disparities in Automatic Speech Recognition TechnologypersoDA: Personalized Data Augmentation for Personalized ASROptimized Self-supervised Training with BEST-RQ for Speech RecognitionCJST: CTC Compressor based Joint Speech and Text Training for
Decoder-Only ASRk2SSL: A Faster and Better Framework for Self-Supervised Speech
Representation LearningCR-CTC: Consistency regularization on CTC for improved speech
recognitionMS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked
Language Modelling methods for learning Speech RepresentationsThe People's Speech: A Large-scale Diverse English Speech Recognition Dataset For Commercial UsageFocused Discriminative Training For Streaming Ctc-trained Automatic Speech Recognition ModelsA Comprehensive Solution To Connect Speech Encoder And Large Language Model For ASRW2v-BERT: Combining Contrastive Learning and Masked Language Modeling
for Self-Supervised Speech Pre-TrainingSqueezeformer: An Efficient Transformer for Automatic Speech RecognitionEfficient conformer: Progressive downsampling and grouped attention for
automatic speech recognitionSLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text
Joint Pre-TrainingTree-constrained Pointer Generator for End-to-end Contextual Speech
RecognitionSelf-supervised Learning with Random-projection Quantizer for Speech
RecognitionZipformer: A faster and better encoder for automatic speech recognitionInjecting Text in Self-Supervised Speech PretrainingExploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical StudyFew-Shot Speaker Identification Using Lightweight Prototypical Network
with Feature Grouping and InteractionImproving Streaming Transformer Based ASR Under a Framework of
Self-supervised LearningPre-Training Transformer Decoder for End-to-End ASR Model with Unpaired
Speech DataEfficient Training of Neural Transducer for Speech RecognitionFine-Tuning Automatic Speech Recognition for People with Parkinson's: An
Effective Strategy for Enhancing Speech Technology AccessibilitySoundChoice: Grapheme-to-Phoneme Models with Semantic DisambiguationRobust Acoustic and Semantic Contextual Biasing in Neural Transducers
for Speech RecognitionOn lattice-free boosted MMI training of HMM and CTC-based full-context
ASR modelsUSTED: Improving ASR with a Unified Speech and Text Encoder-DecoderUnsupervised Data Selection via Discrete Speech Representation for ASRAdaptable End-to-End ASR Models using Replaceable Internal LMs and
Residual SoftmaxAVFormer: Injecting Vision into Frozen Speech Models for Zero-Shot
AV-ASRVALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text
to Speech SynthesizersWav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs
for Robust Speech RecognitionHuBERT-EE: Early Exiting HuBERT for Efficient Speech RecognitionJoint Encoder-Decoder Self-Supervised Pre-training for ASRpMCT: Patched Multi-Condition Training for Robust Speech RecognitionThe People's Speech: A Large-Scale Diverse English Speech Recognition
Dataset for Commercial UsageGlobal Normalization for Streaming Speech Recognition in a Modular
FrameworkE-Branchformer: Branchformer with Enhanced merging for speech
recognitionFAT-HuBERT: Front-end Adaptive Training of Hidden-unit BERT for
Distortion-Invariant Robust Speech RecognitionSA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASRPerformance-Efficiency Trade-offs in Unsupervised Pre-training for
Speech RecognitionAudiobox: Unified Audio Generation with Natural Language PromptsAdvancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces
and ConformersImproving Accent Identification and Accented Speech Recognition Under a
Framework of Self-supervised LearningOn Language Model Integration for RNN Transducer based Speech
RecognitionIntroducing ECAPA-TDNN and Wav2Vec2.0 Embeddings to Stuttering DetectionAnalysis of Self-Attention Head Diversity for Conformer-based Automatic
Speech RecognitionAn Embarrassingly Simple Approach for LLM with Strong ASR CapacityImproving Pseudo-label Training For End-to-end Speech Recognition Using
Gradient MaskLocality Matters: A Locality-Biased Linear Attention for Automatic
Speech RecognitionAuditory-Based Data Augmentation for End-to-End Automatic Speech
RecognitionWav2Vec-Aug: Improved self-supervised training with limited dataSpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder
Based Speech-Text Pre-trainingTESSP: Text-Enhanced Self-Supervised Speech Pre-trainingSvarah: Evaluating English ASR Systems on Indian AccentsLow-rank Adaptation of Large Language Model Rescoring for
Parameter-Efficient Speech RecognitionSSAST: Self-Supervised Audio Spectrogram TransformerConsistent Training and Decoding For End-to-end Speech Recognition Using
Lattice-free MMISelf-Supervised Learning for speech recognition with Intermediate layer
supervisionMLP-ASR: Sequence-length agnostic all-MLP architectures for speech
recognitionRepresentative Subset Selection for Efficient Fine-Tuning in
Self-Supervised Speech RecognitionDeploying self-supervised learning in the wild for hybrid automatic
speech recognitionSpeaker Identification using Speech RecognitionUconv-Conformer: High Reduction of Input Sequence Length for End-to-End
Speech RecognitionSpeeChain: A Speech Toolkit for Large-Scale Machine Speech ChainProgressive unsupervised domain adaptation for ASR using ensemble models
and multi-stage trainingA Cross-Modal Approach to Silent Speech with LLM-Enhanced RecognitionImproving Confidence Estimation on Out-of-Domain Data for End-to-End
Speech RecognitionKnowledge Distillation for Neural Transducers from Large Self-Supervised
Pre-trained ModelsASR4REAL: An extended benchmark for speech modelsImproving Noise Robustness of Contrastive Speech Representation Learning
with Speech ReconstructionDifferentially Private Speaker AnonymizationAsk2Mask: Guided Data Selection for Masked Speech ModelingTransformer-based Streaming ASR with Cumulative AttentionSpectral Modification Based Data Augmentation For Improving End-to-End
ASR For Children's SpeechSelf-Supervised Speech Representations Preserve Speech Characteristics
while Anonymizing VoicesSpeech Sequence Embeddings using Nearest Neighbors Contrastive LearningImproving Streaming End-to-End ASR on Transformer-based Causal Models
with Encoder States Revision StrategiesSpeaker Anonymization with Phonetic Intermediate RepresentationsComparison of Soft and Hard Target RNN-T Distillation for Large-scale
ASRAcoustic-aware Non-autoregressive Spell Correction with Mask Sample
DecodingHMM vs. CTC for Automatic Speech Recognition: Comparison Based on
Full-Sum Training from ScratchStructured State Space Decoder for Speech Recognition and SynthesisInter-KD: Intermediate Knowledge Distillation for CTC-Based Automatic
Speech RecognitionContinual Learning for On-Device Speech Recognition using Disentangled
ConformersFast Entropy-Based Methods of Word-Level Confidence Estimation for
End-To-End Automatic Speech RecognitionUsing External Off-Policy Speech-To-Text Mappings in Contextual
End-To-End Automated Speech RecognitionStructured Pruning of Self-Supervised Pre-trained Models for Speech
Recognition and UnderstandingTime-frequency Network for Robust Speaker RecognitionDistillW2V2: A Small and Streaming Wav2vec 2.0 Based ASR ModelMulti-Head State Space Model for Speech RecognitionHyperConformer: Multi-head HyperMixer for Efficient Speech RecognitionGraph Neural Networks for Contextual ASR with the Tree-Constrained
Pointer GeneratorAdaptive Contextual Biasing for Transducer Based Streaming Speech
RecognitionTowards Selection of Text-to-speech Data to Augment ASR TrainingDecoder-only Architecture for Speech Recognition with CTC Prompts and
Text Data AugmentationLearning from Flawed Data: Weakly Supervised Automatic Speech
RecognitionMulti-resolution HuBERT: Multi-resolution Speech Self-Supervised
Learning with Masked Unit PredictionDiscriminative Speech Recognition Rescoring with Pre-trained Language
ModelsEnd-to-end Multichannel Speaker-Attributed ASR: Speaker Guided Decoder
and Input Feature AnalysisRevisiting the Entropy Semiring for Neural Speech RecognitionImproving ASR Contextual Biasing with Guided AttentionREBORN: Reinforcement-Learned Boundary Segmentation with Iterative
Training for Unsupervised ASRTowards Decoupling Frontend Enhancement and Backend Recognition in
Monaural Robust ASRSpeech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal
LLMs by Teaching the Flow of TimeSOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation
for Low Resource ASRTraining Large ASR Encoders with Differential PrivacyDual Causal/Non-Causal Self-Attention for Streaming End-to-End Speech
RecognitionRelaxed Attention: A Simple Method to Boost Performance of End-to-End
Automatic Speech RecognitionConformer-based End-to-end Speech Recognition With Rotary Position
EmbeddingA Multi-level Acoustic Feature Extraction Framework for Transformer
Based End-to-End Speech RecognitionTask-aware Warping Factors in Mask-based Speech EnhancementASR Rescoring and Confidence Estimation with ELECTRACognitive Coding of SpeechSRU++: Pioneering Fast Recurrence with Attention for Speech RecognitionWord Order Does Not Matter For Speech RecognitionA Unified Speaker Adaptation Approach for ASREfficient Sequence Training of Attention Models using Approximative
RecombinationAutomatic Learning of Subword Dependent Model Scales