WHAMR
Emerging23papers using it
2021first seen
The 'WHAMR!' dataset/benchmark contains a collection of mixed speech signals designed to evaluate speech separation algorithms in challenging acoustic environments with overlapping speakers, background noise, and reverberation.
Papers using WHAMR (23)
- Magnitude-Phase Dual-Path Speech Enhancement Network based on
Self-Supervised Embedding and Perceptual Contrast Stretch BoostingAsymmetric Encoder-Decoder Based on Time-Frequency Correlation for Speech SeparationMoving Speaker Separation via Parallel Spectral-Spatial ProcessingMC-LExt: Multi-Channel Target Speaker Extraction with Onset-Prompted Speaker Conditioning MechanismReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility ScoringListen to Extract: Onset-Prompted Target Speaker ExtractionMossFormer: Pushing the Performance Limit of Monaural Speech Separation
using Gated Single-Head Transformer with Convolution-Augmented Joint
Self-AttentionsConvolutive Prediction for Monaural Speech Dereverberation and
Noisy-Reverberant Speaker SeparationUtterance Weighted Multi-Dilation Temporal Convolutional Networks for
Monaural Speech DereverberationReceptive Field Analysis of Temporal Convolutional Networks for Monaural
Speech DereverberationOn Data Sampling Strategies for Training Neural Network Speech
Separation ModelsOn Time Domain Conformer Models for Monaural Speech Separation in Noisy
Reverberant Acoustic EnvironmentsTF-GridNet: Integrating Full- and Sub-Band Modeling for Speech
SeparationExploring Self-Attention Mechanisms for Speech SeparationA two-stage speaker extraction algorithm under adverse acoustic
conditions using a single-microphoneExploring the Integration of Speech Separation and Recognition with
Self-Supervised Learning RepresentationMossFormer2: Combining Transformer and RNN-Free Recurrent Network for
Enhanced Time-Domain Monaural Speech SeparationBSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech
Enhancement Network based on Self-Supervised EmbeddingUSEF-TSE: Universal Speaker Embedding Free Target Speaker ExtractionDeformable Temporal Convolutional Networks for Monaural Noisy
Reverberant Speech SeparationLibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant
Multi-Talker Speech Separation, ASR and Speaker DiarizationX-CrossNet: A complex spectral mapping approach to target speaker
extraction with cross attention speaker embedding fusionStepwise-Refining Speech Separation Network via Fine-Grained Encoding in
High-order Latent Domain