CSJ
Emerging8papers using it
2022first seen
Corpus of Spontaneous Japanese, or CSJ, is a large-scale database of spontaneous Japanese. It contains speech signal and transcription of about 7 million words along with various annotations like POS and phonetic labels. After describing its design issues, preliminary evaluation of the CSJ was presented. The results su
Papers using CSJ (8)
- Retrieval-Augmented Speech Recognition Approach for Domain ChallengesSpiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early ExitingWhale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech dataStructured State Space Decoder for Speech Recognition and SynthesisA Lexical-aware Non-autoregressive Transformer-based ASR ModelBenchmarking Japanese Speech Recognition on ASR-LLM Setups with
Multi-Pass Augmented Generative Error CorrectionJoint Optimization of Streaming and Non-Streaming Automatic Speech
Recognition with Multi-Decoder and Knowledge DistillationEfficient and Robust Long-Form Speech Recognition with Hybrid
H3-Conformer