← all datasets

GigaSpeech

Canonical
13papers using it
2022first seen

Dataset Card for Gigaspeech Dataset Description GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous sp

Papers using GigaSpeech (13)

GigaSpeech dataset β€” papers, benchmarks & downloads Β· Speech Audio