Cala: Complementary Association Learning For Augmenting Composed Image Retrieval
2024 Β· Xintong Jiang, Yaxiong Wang, Mengjian Li, et al.
Abstract
Composed Image Retrieval (CIR) involves searching for target images based on an image-text pair query. While current methods treat this as a query-target matching problem, we argue that CIR triplets contain additional associations beyond this primary relation. In our paper, we identify two new relations within triplets, treating each triplet as a graph node. Firstly, we introduce the concept of text-bridged image alignment, where the query text serves as a bridge between the query image and the target image. We propose a hinge-based cross-attention mechanism to incorporate this relation into network learning. Secondly, we explore complementary text reasoning, considering CIR as a form of cross-modal retrieval where two images compose to reason about complementary text. To integrate these perspectives effectively, we design a twin attention-based compositor. By combining these complementary associations with the explicit query pair-target image relation, we establish a comprehensive set
Authors
(none)
Tags
Stats
Related papers
- NCL-CIR: Noise-aware Contrastive Learning For Composed Image Retrieval (2025)2.26
- Context-cir: Learning From Concepts In Text For Composed Image Retrieval (2025)4.67
- HINT: Composed Image Retrieval With Dual-path Compositional Contextualized Network (2026)0.78
- From Mapping To Composing: A Two-stage Framework For Zero-shot Composed Image Retrieval (2025)0.00
- CSMCIR: Cot-enhanced Symmetric Alignment With Memory Bank For Composed Image Retrieval (2026)0.00
- Zero-shot Composed Text-image Retrieval (2023)0.00
- Pseudo-triplet Guided Few-shot Composed Image Retrieval (2024)0.00
- Automatic Synthesis Of High-quality Triplet Data For Composed Image Retrieval (2025)0.00