Utterance Weighted Multi-dilation Temporal Convolutional Networks For Monaural Speech Dereverberation
2022 Β· William Ravenscroft, Stefan Goetze, Thomas Hain
Abstract
Speech dereverberation is an important stage in many speech technology applications. Recent work in this area has been dominated by deep neural network models. Temporal convolutional networks (TCNs) are deep learning models that have been proposed for sequence modelling in the task of dereverberating speech. In this work a weighted multi-dilation depthwise-separable convolution is proposed to replace standard depthwise-separable convolutions in TCN models. This proposed convolution enables the TCN to dynamically focus on more or less local information in its receptive field at each convolutional block in the network. It is shown that this weighted multi-dilation temporal convolutional network (WD-TCN) consistently outperforms the TCN across various model configurations and using the WD-TCN model is a more parameter efficient method to improve the performance of the model than increasing the number of convolutional blocks. The best performance improvement over the baseline TCN is 0.55 d
Authors
(none)
Tags
Stats
Related papers
- Receptive Field Analysis Of Temporal Convolutional Networks For Monaural Speech Dereverberation (2022)6.34
- Deformable Temporal Convolutional Networks For Monaural Noisy Reverberant Speech Separation (2022)8.09
- Monaural Speech Enhancement Using A Multi-branch Temporal Convolutional Network (2019)3.58
- Furcanext: End-to-end Monaural Speech Separation With Dynamic Gated Dilated Temporal Convolutional Networks (2019)12.40
- Tecanet: Temporal-contextual Attention Network For Environment-aware Speech Dereverberation (2021)7.50
- TFCN: Temporal-frequential Convolutional Network For Single-channel Speech Enhancement (2022)0.00
- Speech Dereverberation Using Fully Convolutional Networks (2018)13.34
- Speech Dereverberation Using Nonnegative Convolutive Transfer Function And Spectro Temporal Modeling (2017)10.48