📊 Datasets — Awesome Speech Audio
Loading datasets…
Loading datasets…
2,502 datasets & benchmarks — 20 canonical foundations plus emerging datasets mined from recent papers. Each links to the papers that use it.
~1,000 hours of read English audiobook speech with transcripts — a standard ASR benchmark.
The AISHELL-1 dataset is a benchmark for evaluating automatic speech recognition (ASR) systems, specifically focusing on their performance in recognizing and correcting named entities.
IEMOCAP is a dataset used for evaluating speech emotion recognition systems, containing labeled speech data that includes various emotional expressions.
FLEURS is a benchmark used to evaluate speech-to-text translation performance across multiple languages, including the assessment of translation quality in the context of the KUTED dataset for Central Kurdish.
SUPERB is a benchmark dataset that contains a variety of speech processing tasks and is used to evaluate the performance of speech foundation models.
MuST-C is a dataset used to evaluate simultaneous speech-to-speech translation across multiple languages.
VoxCeleb-1 is a dataset used to evaluate speaker verification systems, containing a diverse collection of speech samples from thousands of speakers.
Mozilla's massively-multilingual, crowd-sourced read-speech corpus for speech recognition.
A phonetically-transcribed read-speech corpus widely used for acoustic-phonetic and ASR research.
VoiceBank-DEMAND is a dataset used to evaluate speech quality by providing diverse audio samples with corresponding perceptual mean opinion scores (MOS).
LibriTTS is a dataset that contains high-quality, diverse speech recordings used to evaluate text-to-speech systems' performance in generating natural-sounding speech.
The AMI Meeting Corpus consists of 100 hours of meeting recordings. The recordings use a range of signals synchronized to a common timeline. These include close-talking and far-field microphones, individual and room-view video cameras, and output from a slide projector and an electronic whiteboard. During the meetings, the participants also have unsynchronized pens available to them that record what is written. The meetings were recorded in English using three different rooms with different acoustic properties, and include mostly non-native speakers. \n
A speaker-recognition dataset of utterances from thousands of celebrities collected from YouTube.
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books. A transcription is provided for each clip. Clips vary in length from 1 to 10 seconds and have a total length of approximately 24 hours.
The CSTR VCTK Corpus includes speech data uttered by 110 English speakers with various accents.
The 'Google Speech Commands' dataset contains a collection of spoken commands used to evaluate keyword spotting performance in speech recognition systems.
Libri-2Mix is a dataset used to evaluate Target Speaker Extraction (TSE) performance in mixed speech scenarios.
The Switchboard dataset is an unlabeled target domain used to evaluate the performance of speech recognition systems, particularly in the context of unsupervised domain adaptation.
UASpeech is a dataset that contains recordings of dysarthric speech used to evaluate the effectiveness of speech recognition systems in handling the acoustic variability associated with dysarthria.
LRS-3 is a dataset used to evaluate audio-visual speech recognition (AVSR) systems, containing diverse spoken conversations with human-annotated transcriptions.
The WSJ0-2Mix dataset/benchmark contains mixed speech signals from the Wall Street Journal corpus and is used to evaluate the performance of speech separation models, particularly in the presence of noisy references.
Dataset Card for the Callhome dataset for speaker diarization The CALLHOME Corpus is a collection of unscripted telephone conversations between native speakers in Chinese, English, German, Japanese and Spanish. This is a processed version of the original Callhome dataset from the TalkBank corpora taken from here. It contains subsets in Chinese, English, German, Japanese and Spanish: More information on the Chinese subset More information on the English subset More information on… See the full description on the dataset page: https://huggingface.co/datasets/talkbank/callhome.
CoVoST-2 is a multilingual speech-to-text translation dataset used to evaluate machine translation systems across multiple languages.
LibriMix is a dataset that contains mixed speech recordings of multiple talkers and is used to evaluate multi-talker automatic speech recognition (MT-ASR) systems.
'LibriCSS' is a dataset used to evaluate speech separation performance in challenging acoustic environments, containing recordings of overlapping speakers and background noise.
The CHiME-4 dataset is used to evaluate automatic speech recognition (ASR) systems in noisy environments, containing recordings of speech mixed with various types of background noise.
~2M 10-second YouTube clips labeled with 600+ audio-event classes.
Dataset Card for "slurp" More Information needed
The 'WHAMR!' dataset/benchmark contains a collection of mixed speech signals designed to evaluate speech separation algorithms in challenging acoustic environments with overlapping speakers, background noise, and reverberation.
The 'AliMeeting' dataset is a large-scale conversational dataset used to evaluate end-to-end speaker diarization and recognition systems.
VoxCeleb2 Dataset This is the VoxCeleb2 dataset, a large-scale speaker identification dataset. Dataset Description VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube. Files vox2_dev_mp4_part*: Multipart archive containing MP4 video files vox2_dev_txt: Text files with speaker/utterance metadata vox2_meta.csv: Dataset metadata Usage To extract the multipart archive: # Using 7zip 7z x… See the full description on the dataset page: https://huggingface.co/datasets/Reverb/voxceleb2.
audiocaps HuggingFace mirror of official data repo.
MELD is a dataset that contains multi-modal emotional dialogues used to evaluate the emotional expressiveness and prosodic features in speech synthesis.
\
TEDLIUM-2 is a dataset used for evaluating automatic speech recognition (ASR) systems, containing transcribed audio recordings of TED Talks.
The 'Mandarin' dataset/benchmark is used to evaluate the effectiveness of speech representation learning methods, particularly in the context of tonal languages, by assessing their ability to handle speaker variation and tone distinctions.
The MSP-Podcast dataset is a benchmark that contains in-the-wild speech data used to evaluate the effectiveness of speech emotion conversion methods.
The TORGO dataset is a benchmark that contains recordings of dysarthric speech used to evaluate the effectiveness of automatic speech recognition (ASR) systems in recognizing and processing speech with abnormal prosody and significant speaker variability.
ASVspoof 2019 is a dataset used to evaluate the effectiveness of algorithms in detecting spoofed speech, specifically focusing on deepfake audio generated by synthetic methods.
2,000 labeled environmental-sound clips across 50 classes.
LRS-2 is a dataset used to evaluate audio-visual speech recognition systems, containing video recordings of speakers pronouncing sentences along with corresponding audio, enabling the assessment of models' performance in recognizing speech from visual cues.
Wenetspeech is a multilingual speech dataset used to evaluate the performance of speech-to-text models in aligning speech and text representations across different languages.
The 'WHAM!' dataset/benchmark contains mixtures of speech and is used to evaluate the performance of noisy speech separation systems.
The 'WSJ' dataset is a benchmark for speech recognition that contains transcribed audio data from the Wall Street Journal, used to evaluate the performance of speech recognition systems.
CHiME-3 is a dataset that contains recordings of speech mixed with background noise at various signal-to-noise ratios, used to evaluate speech quality assessment systems.
Deep Noise Suppression (DNS) Challenge - Interspeech 2020 This repository contains the datasets and scripts required for the DNS challenge. For more details about the challenge, please visit https://dns-challenge.azurewebsites.net/ and refer to our paper. Repo details: The datasets directory contains the clean speech and noise clips. The NSNet-baseline directory contains the inference scripts and the ONNX model for the baseline Speech Enhancer called Noise Suppression… See the full description on the dataset page: https://huggingface.co/datasets/ltnghia/DNS-Challenge.
Dataset Card for Gigaspeech Dataset Description GigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training. The transcribed audio data is collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. Example Usage The training split has several configurations of… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech.
The 'In-the-Wild' dataset/benchmark contains real-world audio recordings used to evaluate the performance of audio deepfake detection methods under variable conditions.
LibriSpeechMix is a dataset used to evaluate multi-speaker automatic speech recognition (ASR) by simulating overlapping conversations in English.
The ML-SUPERB dataset is a benchmark used to evaluate multilingual automatic speech recognition (ASR) performance, specifically focusing on the effectiveness of discrete token representations in improving model accuracy.
The 'Speechocean-762' dataset contains 5,000 utterances used to evaluate L2 English pronunciation across multiple aspects, including accuracy, fluency, prosody, and completeness.
The AISHELL-2 dataset is a Mandarin speech recognition benchmark used to evaluate automatic speech recognition (ASR) systems, particularly in the context of streaming recognition.
The 'ASVspoof 5' dataset is a benchmark used to evaluate the effectiveness of speech deepfake detection systems by providing a collection of spoofed and genuine speech samples.
CREMAD is a speech emotion recognition dataset used to evaluate the performance of models in identifying emotions from speech.
The 'Hindi' dataset/benchmark contains transcriptions in Hindi and is used to evaluate the performance of Automatic Speech Recognition (ASR) systems through fine-grained Part-of-Speech (PoS)-wise error characterization and alignment analysis.