Speaker Recognition Based On Deep Learning: An Overview
2020 Β· Zhongxin Bai, Xiao-Lei Zhang
Abstract
Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progress. In this paper, we review several major subtasks of speaker recognition, including speaker verification, identification, diarization, and robust speaker recognition, with a focus on deep-learning-based methods. Because the major advantage of deep learning over conventional methods is its representation ability, which is able to produce highly abstract embedding features from utterances, we first pay close attention to deep-learning-based speaker feature extraction, including the inputs, network structures, temporal pooling strategies, and objective functions respectively, which are the fundamental components of many speaker recognition subtasks. Then, we make an overview of speaker diarization, with an emphasis of recent supervised, end-to-end, and online diarizatio
Authors
(none)
Tags
Stats
Related papers
- Deep Learning Methods In Speaker Recognition: A Review (2019)10.35
- Overview Of Speaker Modeling And Its Applications: From The Lens Of Deep Speaker Representation Learning (2024)10.74
- Supervised Speech Separation Based On Deep Learning: An Overview (2017)0.00
- Adversarial Attack And Defense Strategies For Deep Speaker Recognition Systems (2020)13.39
- Deep Speaker Embeddings For Far-field Speaker Recognition On Short Utterances (2020)11.29
- An Overview Of Deep-learning-based Audio-visual Speech Enhancement And Separation (2020)18.31
- An Overview Of Voice Conversion And Its Challenges: From Statistical Modeling To Deep Learning (2020)18.53
- Speaker Diarization Using Deep Recurrent Convolutional Neural Networks For Speaker Embeddings (2017)9.41