Disentangled-transformer: An Explainable End-to-end Automatic Speech Recognition Model With Speech Content-context Separation

·2024

arXiv:wang2024disentangled ↗Google Scholar ↗Semantic Scholar ↗

Speech Recognition Text-to-Speech Speech Translation

Abstract

End-to-end transformer-based automatic speech recognition (ASR) systems often capture multiple speech traits in their learned representations that are highly entangled, leading to a lack of interpretability. In this study, we propose the explainable Disentangled-Transformer, which disentangles the internal representations into sub-embeddings with explicit content and speaker traits based on varying temporal resolutions. Experimental results show that the proposed Disentangled-Transformer produces a clear speaker identity, separated from the speech content, for speaker diarization while improving ASR performance.

Abstract

Related papers