The Ustc-ximalaya System For The ICASSP 2022 Multi-channel Multi-party Meeting Transcription (m2met) Challenge
2022 Β· Maokui He, Xiang Lv, Weilin Zhou, et al.
Abstract
We propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-Channel Multi-Party Meeting Transcription (M2MeT) challenge. These techniques are designed to handle multi-speaker conversations in real-world meeting scenarios with high speaker-overlap ratios and under heavy reverberant and noisy condition. First, for data preparation and augmentation in training TS-VAD models, speech data containing both real meetings and simulated indoor conversations are used. Second, in refining results obtained after TS-VAD based decoding, we perform a series of post-processing steps to improve the VAD results needed to reduce diarization error rates (DERs). Tested on the ALIMEETING corpus, the newly released Mandarin meeting dataset used in M2MeT, we demonstrate that our proposed system can decrease the DER by up to 66.55/60.59% relatively when compared with classical clustering based
Authors
(none)
Tags
Stats
Related papers
- The CUHK-TENCENT Speaker Diarization System For The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (2022)7.81
- The Volcspeech System For The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (2022)5.84
- Cross-channel Attention-based Target Speaker Voice Activity Detection: Experimental Results For M2met Challenge (2022)10.07
- Summary On The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (2022)10.35
- Royalflush Speaker Diarization System For ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (2022)0.00
- The Xmuspeech System For Multi-channel Multi-party Meeting Transcription Challenge (2022)0.00
- The Second Multi-channel Multi-party Meeting Transcription Challenge (m2met) 2.0): A Benchmark For Speaker-attributed ASR (2023)6.77
- Target-speaker Voice Activity Detection With Improved I-vector Estimation For Unknown Number Of Speaker (2021)10.97