End-to-end Model For Speech Enhancement By Consistent Spectrogram Masking
2019 Β· Xingjian Du, Mengyao Zhu, Xuan Shi, et al.
Abstract
Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT) spectrogram based training targets, e.g. Complex RatioMask (cRM) [1]. However, masking on spectrogram would violentits consistency constraints. In this work, we prove that theinconsistent problem enlarges the solution space of the speechenhancement model and causes unintended artifacts. ConsistencySpectrogram Masking (CSM) is proposed to estimate the complexspectrogram of a signal with the consistency constraint in asimple but not trivial way. The experiments comparing ourCSM based end-to-end model with other methods are conductedto confirm that the CSM accelerate the model training andhave significant improvements in speech quality. From ourexperimental results, we assured that our method could enha
Authors
(none)
Tags
Stats
Related papers
- Magnitude-and-phase-aware Speech Enhancement With Parallel Sequence Modeling (2023)3.58
- Phase Aware Speech Enhancement Using Realisation Of Complex-valued LSTM (2020)0.00
- Phase-incorporating Speech Enhancement Based On Complex-valued Gaussian Process Latent Variable Model (2016)0.00
- Phase-aware Speech Enhancement With Deep Complex U-net (2019)0.00
- An Explicit Consistency-preserving Loss Function For Phase Reconstruction And Speech Enhancement (2024)2.26
- Explicit Estimation Of Magnitude And Phase Spectra In Parallel For High-quality Speech Enhancement (2023)11.19
- Spectral Oversubtraction? An Approach For Speech Enhancement After Robot Ego Speech Filtering In Semi-real-time (2024)0.00
- Complex Spectral Mapping With Attention Based Convolution Recurrent Neural Network For Speech Enhancement (2021)0.00