WSJ-0-2mix
Emerging30papers using it
2021first seen
The WSJ0-2Mix dataset/benchmark contains mixed speech signals from the Wall Street Journal corpus and is used to evaluate the performance of speech separation models, particularly in the presence of noisy references.
Papers using WSJ-0-2mix (30)
- EDSep: An Effective Diffusion-Based Method for Speech Source SeparationA Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy ReferencesDynamic Slimmable Networks for Efficient Speech SeparationListen to Extract: Onset-Prompted Target Speaker ExtractionAn Investigation on Speaker Augmentation for End-to-End Speaker ExtractionESPnet-SE++: Speech Enhancement for Robust Speech Recognition,
Translation, and UnderstandingX-SepFormer: End-to-end Speaker Extraction Network with Explicit
Optimization on Speaker ConfusionDiscretization and Re-synthesis: an alternative method to solve the
Cocktail Party ProblemResource-Efficient Separation TransformerDual-path Mamba: Short and Long-term Bidirectional Selective Structured
State Space Models for Speech SeparationAmbiSep: Ambisonic-to-Ambisonic Reverberant Speech Separation Using
Transformer NetworksAn Exploration of Self-Supervised Pretrained Representations for
End-to-End Speech RecognitionSPMamba: State-space model is all you need in speech separationTF-GridNet: Making Time-Frequency Domain Models Great Again for Monaural
Speaker SeparationOn Time Domain Conformer Models for Monaural Speech Separation in Noisy
Reverberant Acoustic EnvironmentsTF-GridNet: Integrating Full- and Sub-Band Modeling for Speech
SeparationExploring Self-Attention Mechanisms for Speech SeparationConditional Diffusion Model for Target Speaker ExtractionSQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASRUX-NET: Filter-and-Process-based Improved U-Net for Real-time
Time-domain Audio SeparationDiffusion-based Generative Speech Source SeparationMulti-Scale Feature Fusion Transformer Network for End-to-End Single
Channel Speech SeparationMulti-Dimensional and Multi-Scale Modeling for Speech Separation
Optimized by Discriminative LearningSpeech Separation based on Contrastive Learning and Deep ModularizationWeakly-Supervised Speech Pre-training: A Case Study on Target Speech
RecognitionUSEF-TSE: Universal Speaker Embedding Free Target Speaker ExtractionSpeech Separation using Neural Audio Codecs with Embedding LossImproving Target Speaker Extraction with Sparse LDA-transformed Speaker
EmbeddingsSPGM: Prioritizing Local Features for enhanced speech separation
performanceX-CrossNet: A complex spectral mapping approach to target speaker
extraction with cross attention speaker embedding fusion