VGGSound
Emerging7papers using it
2022first seen
VGGSound VGG-Sound is an audio-visual correspondent dataset consisting of short clips of audio sounds, extracted from videos uploaded to YouTube. Homepage: https://www.robots.ox.ac.uk/~vgg/data/vggsound/ Paper: https://arxiv.org/abs/2004.14368 Github: https://github.com/hche11/VGGSound Analysis 310+ classes: VGG-Sound
Papers using VGGSound (7)
- FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic PrecisionDiffGAP: A Lightweight Diffusion Module in Contrastive Space for
Bridging Cross-Model GapGMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative PretrainingVGGSounder: Audio-Visual Evaluations for Foundation ModelsTraining-free Multimodal Guidance For Video To Audio GenerationONE-PEACE: Exploring One General Representation Model Toward Unlimited
ModalitiesMultimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models