Multi-task Deep Residual Echo Suppression With Echo-aware Loss
2022 Β· Shimin Zhang, Ziteng Wang, Jiayao Sun, et al.
Abstract
This paper introduces the NWPU Team's entry to the ICASSP 2022 AEC Challenge. We take a hybrid approach that cascades a linear AEC with a neural post-filter. The former is used to deal with the linear echo components while the latter suppresses the residual non-linear echo components. We use gated convolutional F-T-LSTM neural network (GFTNN) as the backbone and shape the post-filter by a multi-task learning (MTL) framework, where a voice activity detection (VAD) module is adopted as an auxiliary task along with echo suppression, with the aim to avoid over suppression that may cause speech distortion. Moreover, we adopt an echo-aware loss function, where the mean square error (MSE) loss can be optimized particularly for every time-frequency bin (TF-bin) according to the signal-to-echo ratio (SER), leading to further suppression on the echo. Extensive ablation study shows that the time delay estimation (TDE) module in neural post-filter leads to better perceptual quality, and an adaptiv
Authors
(none)
Tags
Stats
Related papers
- Neuralecho: A Self-attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement (2022)0.00
- Joint Echo Cancellation And Noise Suppression Based On Cascaded Magnitude And Complex Mask Estimation (2021)0.00
- An Exploration Of Task-decoupling On Two-stage Neural Post Filter For Real-time Personalized Acoustic Echo Cancellation (2023)0.00
- Deepvqe: Real Time Deep Voice Quality Enhancement For Joint Acoustic Echo Cancellation, Noise Suppression And Dereverberation (2023)0.00
- Deep Residual Echo Suppression And Noise Reduction: A Multi-input FCRN Approach In A Hybrid Speech Enhancement System (2021)8.09
- F-T-LSTM Based Complex Network For Joint Acoustic Echo Cancellation And Speech Enhancement (2021)11.19
- Acoustic Echo Cancellation With The Dual-signal Transformation LSTM Network (2020)12.93
- Implicit Acoustic Echo Cancellation For Keyword Spotting And Device-directed Speech Detection (2021)3.58