You'll Never Walk Alone: A Sketch And Text Duet For Fine-grained Image Retrieval
2024 Β· Subhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, et al.
Abstract
Two primary input modalities prevail in image retrieval: sketch and text. While text is widely used for inter-category retrieval tasks, sketches have been established as the sole preferred modality for fine-grained image retrieval due to their ability to capture intricate visual details. In this paper, we question the reliance on sketches alone for fine-grained image retrieval by simultaneously exploring the fine-grained representation capabilities of both sketch and text, orchestrating a duet between the two. The end result enables precise retrievals previously unattainable, allowing users to pose ever-finer queries and incorporate attributes like colour and contextual cues from text. For this purpose, we introduce a novel compositionality framework, effectively combining sketches and text using pre-trained CLIP models, while eliminating the need for extensive fine-grained textual descriptions. Last but not least, our system extends to novel applications in composed image retrieval, d
Authors
(none)
Tags
Stats
Related papers
- Sketch And Text Synergy: Fusing Structural Contours And Descriptive Attributes For Fine-grained Image Retrieval (2026)0.00
- A Sketch Is Worth A Thousand Words: Image Retrieval With Text And Sketch (2022)10.35
- Learning Cross-modal Deep Embeddings For Multi-object Image Retrieval Using Text And Sketch (2018)9.59
- Cross-modal Hierarchical Modelling For Fine-grained Sketch Based Image Retrieval (2020)6.77
- Sketching Without Worrying: Noise-tolerant Sketch-based Image Retrieval (2022)12.74
- Sketch Less For More: On-the-fly Fine-grained Sketch Based Image Retrieval (2020)15.28
- Freeview Sketching: View-aware Fine-grained Sketch-based Image Retrieval (2024)6.34
- Cross-modal Fusion Distillation For Fine-grained Sketch-based Image Retrieval (2022)2.68