Aishell-4
Emerging9papers using it
2021first seen
AISHELL-4 is a dataset used to evaluate speaker diarization and recognition systems, containing conversational audio data that facilitates the assessment of speaker classification performance.
Papers using Aishell-4 (9)
- SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech TranscriptionBalancing ASR and diarization in end-to-end LLMs for multi-talker speech recognitionSpeaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASRJoint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and RecognitionLightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic TransformExploiting Single-Channel Speech For Multi-channel End-to-end Speech
RecognitionExploiting Single-Channel Speech for Multi-Channel End-to-End Speech
Recognition: A Comparative StudyToken-level Speaker Change Detection Using Speaker Difference and Speech
Content via Continuous Integrate-and-fireASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in
Meetings