Semantic Residual For Multimodal Unified Discrete Representation
2024 Β· Hai Huang, Shulei Wang, Yan Xia
Abstract
Recent research in the domain of multimodal unified representations predominantly employs codebook as representation forms, utilizing Vector Quantization(VQ) for quantization, yet there has been insufficient exploration of other quantization representation forms. Our work explores more precise quantization methods and introduces a new framework, Semantic Residual Cross-modal Information Disentanglement (SRCID), inspired by the numerical residual concept inherent to Residual Vector Quantization (RVQ). SRCID employs semantic residual-based information disentanglement for multimodal data to better handle the inherent discrepancies between different modalities. Our method enhances the capabilities of unified multimodal representations and demonstrates exceptional performance in cross-modal generalization and cross-modal zero-shot retrieval. Its average results significantly surpass existing state-of-the-art models, as well as previous attempts with RVQ and Finite Scalar Quantization (FSQ)
Authors
(none)
Tags
Stats
Related papers
- Cross-modal Discrete Representation Learning (2021)10.61
- Universal Vision-language Dense Retrieval: Learning A Unified Representation Space For Multi-modal Retrieval (2022)3.45
- Dynamic Visual Semantic Sub-embeddings And Fast Re-ranking (2023)0.00
- Unified Representation Learning For Cross Model Compatibility (2020)5.24
- Preserving Semantic Neighborhoods For Robust Cross-modal Retrieval (2020)10.07
- Semcore: A Semantic-enhanced Generative Cross-modal Retrieval Framework With Mllms (2025)0.00
- Combating Visual Neglect And Semantic Drift In Large Multimodal Models For Enhanced Cross-modal Retrieval (2026)0.00
- Generalized Multi-view Embedding For Visual Recognition And Cross-modal Retrieval (2016)14.69