← all datasets

WavCaps

Emerging
3papers using it
5,704HF downloads
55HF likes
2024first seen

WavCaps WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset). Paper: https://arxiv.org/ab

Papers using WavCaps (3)

WavCaps dataset β€” papers, benchmarks & downloads Β· Multimodal