ALADIN: Distilling Fine-grained Alignment Scores For Efficient Image-text Matching And Retrieval
2022 Β· Nicola Messina, Matteo Stefanini, Marcella Cornia, et al.
Abstract
Image-text matching is gaining a leading role among tasks involving the joint understanding of vision and language. In literature, this task is often used as a pre-training objective to forge architectures able to jointly deal with images and texts. Nonetheless, it has a direct downstream application: cross-modal retrieval, which consists in finding images related to a given query text or vice-versa. Solving this task is of critical importance in cross-modal search engines. Many recent methods proposed effective solutions to the image-text matching problem, mostly using recent large vision-language (VL) Transformer networks. However, these models are often computationally expensive, especially at inference time. This prevents their adoption in large-scale cross-modal retrieval scenarios, where results should be provided to the user almost instantaneously. In this paper, we propose to fill in the gap between effectiveness and efficiency by proposing an ALign And DIstill Network (ALADIN)
Authors
(none)
Tags
Stats
Related papers
- A New Fine-grained Alignment Method For Image-text Matching (2023)0.00
- Towards Fast And Accurate Image-text Retrieval With Self-supervised Fine-grained Alignment (2023)11.99
- MCAD: Multi-teacher Cross-modal Alignment Distillation For Efficient Image-text Retrieval (2023)3.58
- Optimizing CLIP Models For Image Retrieval With Maintained Joint-embedding Alignment (2024)6.34
- Transformer Reasoning Network For Image-text Matching And Retrieval (2020)16.15
- Fine-grained Visual Textual Alignment For Cross-modal Retrieval Using Transformer Encoders (2020)19.48
- Modest-align: Data-efficient Alignment For Vision-language Models (2025)0.00
- Efficient Medical Vision-language Alignment Through Adapting Masked Vision Models (2025)5.74