End-to-end Far-field Speech Recognition With Unified Dereverberation And Beamforming
2020 Β· Wangyou Zhang, Aswin Shanmugam Subramanian, Xuankai Chang, et al.
Abstract
Despite successful applications of end-to-end approaches in multi-channel speech recognition, the performance still degrades severely when the speech is corrupted by reverberation. In this paper, we integrate the dereverberation module into the end-to-end multi-channel speech recognition system and explore two different frontend architectures. First, a multi-source mask-based weighted prediction error (WPE) module is incorporated in the frontend for dereverberation. Second, another novel frontend architecture is proposed, which extends the weighted power minimization distortionless response (WPD) convolutional beamformer to perform simultaneous separation and dereverberation. We derive a new formulation from the original WPD, which can handle multi-source input, and replace eigenvalue decomposition with the matrix inverse operation to make the back-propagation algorithm more stable. The above two architectures are optimized in a fully end-to-end manner, only using the speech recognitio
Authors
(none)
Tags
Stats
Related papers
- End-to-end Dereverberation, Beamforming, And Speech Recognition With Improved Numerical Stability And Advanced Frontend (2021)10.97
- A Unified Convolutional Beamformer For Simultaneous Denoising And Dereverberation (2018)14.15
- WPD++: An Improved Neural Beamformer For Simultaneous Speech Separation And Dereverberation (2020)6.77
- Joint Multi-channel Dereverberation And Noise Reduction Using A Unified Convolutional Beamformer With Sparse Priors (2021)0.00
- A Unified Multichannel Far-field Speech Recognition System: Combining Neural Beamforming With Attention Based End-to-end Model (2024)0.00
- Task-specific Optimization Of Virtual Channel Linear Prediction-based Speech Dereverberation Front-end For Far-field Speaker Verification (2021)2.26
- End-to-end Integration Of Speech Recognition, Dereverberation, Beamforming, And Self-supervised Learning Representation (2022)8.60
- Run-time Adaptation Of Neural Beamforming For Robust Speech Dereverberation And Denoising (2024)0.00