← all datasets

CSJ

Emerging
8papers using it
2022first seen

Corpus of Spontaneous Japanese, or CSJ, is a large-scale database of spontaneous Japanese. It contains speech signal and transcription of about 7 million words along with various annotations like POS and phonetic labels. After describing its design issues, preliminary evaluation of the CSJ was presented. The results su

Papers using CSJ (8)

CSJ dataset β€” papers, benchmarks & downloads Β· Speech Audio