Complex Spectral Mapping With Attention Based Convolution Recurrent Neural Network For Speech Enhancement
2021 Β· Liming Zhou, Yongyu Gao, Ziluo Wang, et al.
Abstract
Speech enhancement has benefited from the success of deep learning in terms of intelligibility and perceptual quality. Conventional time-frequency (TF) domain methods focus on predicting TF-masks or speech spectrum,via a naive convolution neural network or recurrent neural network.Some recent studies were based on Complex spectral Mapping convolution recurrent neural network (CRN) . These models skiped directly from encoder layers' output and decoder layers' input ,which maybe thoughtless. We proposed an attention mechanism based skip connection between encoder and decoder layers,namely Complex Spectral Mapping With Attention Based Convolution Recurrent Neural Network (CARN).Compared with CRN model,the proposed CARN model improved more than 10% relatively at several metrics such as PESQ,CBAK,COVL,CSIG and son,and outperformed the place 1st model in both real time and non-real time track of the DNS Challenge 2020 at these metrics.
Authors
(none)
Tags
Stats
Related papers
- DCCRN: Deep Complex Convolution Recurrent Network For Phase-aware Speech Enhancement (2020)20.78
- DCCRN+: Channel-wise Subband DCCRN With SNR Estimation For Speech Enhancement (2021)0.00
- Single Channel Speech Enhancement Using Temporal Convolutional Recurrent Neural Networks (2020)5.84
- Real-time Monaural Speech Enhancement With Short-time Discrete Cosine Transform (2021)0.00
- Multi-loss Convolutional Network With Time-frequency Attention For Speech Enhancement (2023)0.00
- Effcrn: An Efficient Convolutional Recurrent Network For High-performance Speech Enhancement (2023)5.84
- A Deep Representation Learning-based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder (2023)7.16
- SICRN: Advancing Speech Enhancement Through State Space Model And Inplace Convolution Techniques (2024)7.81