Spectral Clustering-aware Learning Of Embeddings For Speaker Diarisation
2022 Β· Evonne P. C. Lee, Guangzhi Sun, Chao Zhang, et al.
Abstract
In speaker diarisation, speaker embedding extraction models often suffer from the mismatch between their training loss functions and the speaker clustering method. In this paper, we propose the method of spectral clustering-aware learning of embeddings (SCALE) to address the mismatch. Specifically, besides an angular prototype cal (AP) loss, SCALE uses a novel affinity matrix loss which directly minimises the error between the affinity matrix estimated from speaker embeddings and the reference. SCALE also includes p-percentile thresholding and Gaussian blur as two important hyper-parameters for spectral clustering in training. Experiments on the AMI dataset showed that speaker embeddings obtained with SCALE achieved over 50% relative speaker error rate reductions using oracle segmentation, and over 30% relative diarisation error rate reductions using automatic segmentation when compared to a strong baseline with the AP-loss-based speaker embeddings.
Authors
(none)
Tags
Stats
Related papers
- Multi-scale Speaker Embedding-based Graph Attention Networks For Speaker Diarisation (2021)8.35
- Improved Large-margin Softmax Loss For Speaker Diarisation (2019)6.34
- Assessing The Robustness Of Spectral Clustering For Deep Speaker Diarization (2024)3.58
- Multi-class Spectral Clustering With Overlaps For Speaker Diarization (2020)10.35
- Speaker Diarisation Using 2D Self-attentive Combination Of Embeddings (2019)9.92
- Deep Self-supervised Hierarchical Clustering For Speaker Diarization (2020)5.24
- Self-tuning Spectral Clustering For Speaker Diarization (2024)3.81
- Discriminative Neural Clustering For Speaker Diarisation (2019)10.07