Overview Of Speaker Modeling And Its Applications: From The Lens Of Deep Speaker Representation Learning
2024 Β· Shuai Wang, Zhengyang Chen, Kong Aik Lee, et al.
Abstract
Speaker individuality information is among the most critical elements within speech signals. By thoroughly and accurately modeling this information, it can be utilized in various intelligent speech applications, such as speaker recognition, speaker diarization, speech synthesis, and target speaker extraction. In this overview, we present a comprehensive review of neural approaches to speaker representation learning from both theoretical and practical perspectives. Theoretically, we discuss speaker encoders ranging from supervised to self-supervised learning algorithms, standalone models to large pretrained models, pure speaker embedding learning to joint optimization with downstream tasks, and efforts toward interpretability. Practically, we systematically examine approaches for robustness and effectiveness, introduce and compare various open-source toolkits in the field. Through the systematic and comprehensive review of the relevant literature, research activities, and resources, we
Authors
(none)
Tags
Stats
Related papers
- Speaker Recognition Based On Deep Learning: An Overview (2020)18.86
- Wespeaker: A Research And Production Oriented Speaker Embedding Learning Toolkit (2022)6.22
- Deep Representation Learning In Speech Processing: Challenges, Recent Advances, And Future Trends (2020)0.00
- Towards Neural Speaker Modeling In Multi-party Conversation: The Task, Dataset, And Models (2017)6.34
- Investigation Of Speaker Representation For Target-speaker Speech Processing (2024)4.52
- How To Improve Your Speaker Embeddings Extractor In Generic Toolkits (2018)9.76
- Investigation Of Speaker-adaptation Methods In Transformer Based ASR (2020)0.00
- Deep Learning Methods In Speaker Recognition: A Review (2019)10.35