Supervised Hierarchical Clustering Using Graph Neural Networks For Speaker Diarization
2023 Β· Prachi Singh, Amrit Kaul, Sriram Ganapathy
Abstract
Conventional methods for speaker diarization involve windowing an audio file into short segments to extract speaker embeddings, followed by an unsupervised clustering of the embeddings. This multi-step approach generates speaker assignments for each segment. In this paper, we propose a novel Supervised HierArchical gRaph Clustering algorithm (SHARC) for speaker diarization where we introduce a hierarchical structure using Graph Neural Network (GNN) to perform supervised clustering. The supervision allows the model to update the representations and directly improve the clustering performance, thus enabling a single-step approach for diarization. In the proposed work, the input segment embeddings are treated as nodes of a graph with the edge weights corresponding to the similarity scores between the nodes. We also propose an approach to jointly update the embedding extractor and the GNN model to perform end-to-end speaker diarization (E2E-SHARC). During inference, the hierarchical cluste
Authors
(none)
Tags
Stats
Related papers
- End-to-end Supervised Hierarchical Graph Clustering For Speaker Diarization (2024)5.24
- Deep Self-supervised Hierarchical Clustering For Speaker Diarization (2020)5.24
- Multi-scale Speaker Embedding-based Graph Attention Networks For Speaker Diarisation (2021)8.35
- Speaker Diarization Using Two-pass Leave-one-out Gaussian PLDA Clustering Of DNN Embeddings (2021)2.26
- Low-latency Online Speaker Diarization With Graph-based Label Generation (2021)8.60
- Learning Deep Representations By Multilayer Bootstrap Networks For Speaker Diarization (2019)0.00
- Speaker Diarization Using Deep Recurrent Convolutional Neural Networks For Speaker Embeddings (2017)9.41
- Tight Integration Of Neural- And Clustering-based Diarization Through Deep Unfolding Of Infinite Gaussian Mixture Model (2022)8.60