CHiME-4
Emerging26papers using it
2021first seen
The CHiME-4 dataset is used to evaluate automatic speech recognition (ASR) systems in noisy environments, containing recordings of speech mixed with various types of background noise.
Papers using CHiME-4 (26)
- FlexIO: Flexible Single- and Multi-Channel Speech Separation and EnhancementFrontend Token Enhancement for Token-Based Speech RecognitionNoisy Disentanglement with Tri-stage Training for Noise-Robust Speech RecognitionMixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASRReducing the Gap Between Pretrained Speech Enhancement and Recognition
Models Using a Real Speech-Trained Bridging ModuleESPnet-SE++: Speech Enhancement for Robust Speech Recognition,
Translation, and UnderstandingSpeaker Reinforcement Using Target Source Extraction for Robust
Automatic Speech RecognitionWav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs
for Robust Speech RecognitionA Conformer Based Acoustic Model for Robust Automatic Speech RecognitionClosing the Gap Between Time-Domain Multi-Channel Speech Enhancement on
Real and Simulation ConditionsMulti-Variant Consistency based Self-supervised Learning for Robust
Automatic Speech RecognitionRobust Data2vec: Noise-robust Speech Representation Learning for ASR by
Combining Regression and Improved Contrastive LearningImproving Noise Robustness of Contrastive Speech Representation Learning
with Speech ReconstructionDual-Path Style Learning for End-to-End Noise-Robust Speech RecognitionEnd-to-End Integration of Speech Recognition, Speech Enhancement, and
Self-Supervised Learning RepresentationExploration of Adapter for Noise Robust Automatic Speech RecognitionTowards Decoupling Frontend Enhancement and Backend Recognition in
Monaural Robust ASRExploiting Single-Channel Speech For Multi-channel End-to-end Speech
RecognitionExploiting Single-Channel Speech for Multi-Channel End-to-End Speech
Recognition: A Comparative StudyMultiple-hypothesis RNN-T Loss for Unsupervised Fine-tuning and
Self-training of Neural TransducerEnd-to-End Integration of Speech Recognition, Dereverberation,
Beamforming, and Self-Supervised Learning RepresentationGradient Remedy for Multi-Task Learning in End-to-End Noise-Robust
Speech RecognitionStatistical Beamformer Exploiting Non-stationarity and Sparsity with
Spatially Constrained ICA for Robust Speech RecognitionctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech EnhancementSelf-Supervised Learning for Multi-Channel Neural TransducerEvolutionary Prompt Design for LLM-Based Post-ASR Error Correction