Deep Factorization For Speech Signal
2018 Β· Lantian Li, Dong Wang, Yixiang Chen, et al.
Abstract
Various informative factors mixed in speech signals, leading to great difficulty when decoding any of the factors. An intuitive idea is to factorize each speech frame into individual informative factors, though it turns out to be highly difficult. Recently, we found that speaker traits, which were assumed to be long-term distributional properties, are actually short-time patterns, and can be learned by a carefully designed deep neural network (DNN). This discovery motivated a cascade deep factorization (CDF) framework that will be presented in this paper. The proposed framework infers speech factors in a sequential way, where factors previously inferred are used as conditional variables when inferring other factors. We will show that this approach can effectively factorize speech signals, and using these factors, the original speech spectrum can be recovered with a high accuracy. This factorization and reconstruction approach provides potential values for many speech processing tasks,
Authors
(none)
Tags
Stats
Related papers
- Deep Generative Factorization For Speech Signal (2020)0.00
- Mixture Factorized Auto-encoder For Unsupervised Hierarchical Deep Factorization Of Speech Signal (2019)0.00
- Content-context Factorized Representations For Automated Speech Recognition (2022)6.34
- On Investigation Of Unsupervised Speech Factorization Based On Normalization Flow (2019)0.00
- Feature Joint-state Posterior Estimation In Factorial Speech Processing Models Using Deep Neural Networks (2017)3.58
- Full-info Training For Deep Speaker Feature Learning (2017)7.16
- Deep Speaker Feature Learning For Text-independent Speaker Verification (2017)12.54
- Self-supervised Neural Factor Analysis For Disentangling Utterance-level Speech Representations (2023)0.00