Aishell-1
Emerging76papers using it
2021first seen
The AISHELL-1 dataset is a benchmark for evaluating automatic speech recognition (ASR) systems, specifically focusing on their performance in recognizing and correcting named entities.
Papers using Aishell-1 (76)
- PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech RecognitionBreaking Through the Spike: Spike Window Decoding for Accelerated and
Precise Automatic Speech RecognitionJSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASRRetrieval-Augmented Self-Taught Reasoning Model with Adaptive Chain-of-Thought for ASR Named Entity CorrectionIKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech RecognitionStreaming Speech Recognition with Decoder-Only Large Language Models and Latency OptimizationEnd-to-end Speech Recognition with similar length speech and textA Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASRObjective Soups: Multilingual Multi-Task Modeling for Speech ProcessingIML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech ProcessingCR-CTC: Consistency regularization on CTC for improved speech
recognitionM2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for
Enhancing WhisperEffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech
Recognition Architecture with High Accuracy and Inference SpeedZipformer: A faster and better encoder for automatic speech recognitionImproving non-autoregressive end-to-end speech recognition with
pre-trained acoustic and language modelsExploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical StudyUnveiling the Potential of LLM-Based ASR on Chinese Open-Source DatasetsImproving Hybrid CTC/Attention End-to-end Speech Recognition with
Pretrained Acoustic and Language ModelMulti-Level Modeling Units for End-to-End Mandarin Speech RecognitionSpeaker-Aware Mixture of Mixtures Training for Weakly Supervised Speaker
ExtractionNextformer: A ConvNeXt Augmented Conformer For End-To-End Speech
RecognitionImproving Mandarin Speech Recogntion with Block-augmented TransformerConsistent Training and Decoding For End-to-end Speech Recognition Using
Lattice-free MMIParaformer: Fast and Accurate Parallel Transformer for
Non-autoregressive End-to-End Speech RecognitionTowards Unified All-Neural Beamforming for Time and Frequency Domain
Speech SeparationUniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL
ModelsEfficientASR: Speech Recognition Network Compression via Attention
Redundancy and Chunk-Level FFN OptimizationDecoupling recognition and transcription in Mandarin ASRNon-autoregressive Transformer with Unified Bidirectional Decoder for
Automatic Speech RecognitionOn the Effectiveness of Pinyin-Character Dual-Decoding for End-to-End
Mandarin Chinese ASRTransformer-based Streaming ASR with Cumulative AttentionImproving CTC-based ASR Models with Gated Interlayer CollaborationLinguistic-Enhanced Transformer with CTC Embedding for Speech
RecognitionConformer-based End-to-end Speech Recognition With Rotary Position
EmbeddingCross-domain Single-channel Speech Enhancement Model with Bi-projection
Fusion Module for Noise-robust ASRFastCorrect 2: Fast Error Correction on Multiple Candidates for
Automatic Speech RecognitionImproving CTC-based speech recognition via knowledge transferring from
pre-trained language modelsShifted Chunk Encoder for Transformer Based Streaming End-to-End ASRCUSIDE: Chunking, Simulating Future Context and Decoding for Streaming
ASRAn Empirical Study of Language Model Integration for Transducer based
Speech RecognitionMemory-Efficient Training of RNN-Transducer with Sampled SoftmaxA CTC Triggered Siamese Network with Spatial-Temporal Dropout for Speech
RecognitionKnowledge Transfer and Distillation from Autoregressive to
Non-Autoregressive Speech RecognitionPSVRF: Learning to restore Pitch-Shifted Voice without referenceA context-aware knowledge transferring strategy for CTC-based ASRSAN: a robust end-to-end ASR model architectureFast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint
CTC/Attention FramesImproving Noisy Student Training on Non-target Domain Data for Automatic
Speech RecognitionSSCFormer: Push the Limit of Chunk-wise Conformer for Streaming ASR
Using Sequentially Sampled Chunks and Chunked Causal ConvolutionKnowledge Transfer from Pre-trained Language Models to Cif-based Speech
Recognizers via Hierarchical DistillationBeyond Universal Transformer: block reusing with adaptor in Transformer
for automatic speech recognitionPyramid Multi-branch Fusion DCNN with Multi-Head Self-Attention for
Mandarin Speech RecognitionSelf-regularised Minimum Latency Training for Streaming
Transformer-based Speech RecognitionA Lexical-aware Non-autoregressive Transformer-based ASR ModelGNCformer Enhanced Self-attention for Automatic Speech RecognitionRethinking Speech Recognition with A Multimodal Perspective via Acoustic
and Semantic Cooperative DecodingEnhancing the Unified Streaming and Non-streaming Model with Contrastive
LearningResearch on an improved Conformer end-to-end Speech Recognition Model
with R-Drop StructureTST: Time-Sparse Transducer for Automatic Speech RecognitionCIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech
RecognitionApproBiVT: Lead ASR Models to Generalize Better Using Approximated
Bias-Variance Tradeoff Guided Early Stopping and Checkpoint AveragingHypR: A comprehensive study for ASR hypothesis revising with a reference
corpusCross-modal Alignment with Optimal Transport for CTC-based ASRHierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention
for CTC-based ASRSkipformer: A Skip-and-Recover Strategy for Efficient Speech RecognitionMulti-Channel Multi-Speaker ASR Using Target Speaker's Solo SegmentStreaming Decoder-Only Automatic Speech Recognition with Discrete Speech
Units: A Pilot StudyCUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based
Streaming ASRHydraFormer: One Encoder For All Subsampling RatesAn Effective Context-Balanced Adaptation Approach for Long-Tailed Speech
RecognitionLarge Language Model Should Understand Pinyin for Chinese ASR Error
CorrectionBridging Speech and Text: Enhancing ASR with Pinyin-to-Character
Pre-training in LLMsDeep CLAS: Deep Contextual Listen, Attend and SpellSample adaptive data augmentation with progressive schedulingImproving Mandarin End-to-End Speech Recognition with Word N-gram
Language ModelUCorrect: An Unsupervised Framework for Automatic Speech Recognition
Error Correction