AISHELL-2
Emerging12papers using it
2021first seen
The AISHELL-2 dataset is a Mandarin speech recognition benchmark used to evaluate automatic speech recognition (ASR) systems, particularly in the context of streaming recognition.
Papers using AISHELL-2 (12)
- Streaming Speech Recognition with Decoder-Only Large Language Models and Latency OptimizationEnd-to-end Speech Recognition with similar length speech and textEffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech
Recognition Architecture with High Accuracy and Inference SpeedConsistent Training and Decoding For End-to-end Speech Recognition Using
Lattice-free MMIParaformer: Fast and Accurate Parallel Transformer for
Non-autoregressive End-to-End Speech RecognitionFew-Shot Speaker Identification Using Depthwise Separable Convolutional
Network with Channel AttentionA context-aware knowledge transferring strategy for CTC-based ASRImproving Noisy Student Training on Non-target Domain Data for Automatic
Speech RecognitionApproBiVT: Lead ASR Models to Generalize Better Using Approximated
Bias-Variance Tradeoff Guided Early Stopping and Checkpoint AveragingStreaming Decoder-Only Automatic Speech Recognition with Discrete Speech
Units: A Pilot StudySemi-supervised Learning for Code-Switching ASR with Large Language
Model FilterImproving Mandarin End-to-End Speech Recognition with Word N-gram
Language Model