← all papers · overview

Closing The Modality Gap Aligns Group-wise Semantics

Abstract

In multimodal learning, CLIP has been recognized as the \textit\{de facto\} method for learning a shared latent space across multiple modalities, placing similar representations close to each other and moving them away from dissimilar ones. Although CLIP-based losses effectively align modalities at the semantic level, the resulting latent spaces often remain only partially shared, revealing a structural mismatch known as the modality gap. While the necessity of addressing this phenomenon remains debated, particularly given its limited impact on instance-wise tasks (e.g., retrieval), we prove that its influence is instead strongly pronounced in group-level tasks (e.g., clustering). To support this claim, we introduce a novel method designed to consistently reduce this discrepancy in two-modal settings, with a straightforward extension to the general -modal case. Through our extensive evaluation, we demonstrate our novel insight: while reducing the gap provides only marginal or inco

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).