Diffgap: A Lightweight Diffusion Module In Contrastive Space For Bridging Cross-model Gap
2025 Β· Shentong Mo, Zehua Chen, Fan Bao, et al.
Abstract
Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignment of text, video, and audio embeddings via a single contrastive loss. However, these methods often overlook the bidirectional interactions and inherent noises present in each modality, which can crucially impact the quality and efficacy of cross-modal integration. To address this limitation, we introduce DiffGAP, a novel approach incorporating a lightweight generative module within the contrastive space. Specifically, our DiffGAP employs a bidirectional diffusion process tailored to bridge the cross-modal gap more effectively. This involves a denoising process on text and video embeddings conditioned on audio embeddings and vice versa, thus facilitating a more nuanced and robust cross-modal interaction. Our experimental results on VGGSound and AudioCaps datasets demons
Authors
(none)
Tags
Stats
Related papers
- Mmdisco: Multi-modal Discriminator-guided Cooperative Diffusion For Joint Audio And Video Generation (2024)1.91
- Extract And Diffuse: Latent Integration For Improved Diffusion-based Speech And Vocal Enhancement (2024)0.00
- Diff-foley: Synchronized Video-to-audio Synthesis With Latent Diffusion Models (2023)0.00
- Av-link: Temporally-aligned Diffusion Features For Cross-modal Audio-video Generation (2024)0.00
- A Simple But Strong Baseline For Sounding Video Generation: Effective Adaptation Of Audio And Video Diffusion Models For Joint Generation (2024)3.58
- Immersediffusion: A Generative Spatial Audio Latent Diffusion Model (2024)0.00
- Contrastive Conditional Latent Diffusion For Audio-visual Segmentation (2023)10.29
- Syncdiff: Diffusion-based Talking Head Synthesis With Bottlenecked Temporal Visual Prior For Improved Synchronization (2025)4.52