FIGROTD: A Friendly-to-handle Dataset For Image Guided Retrieval With Optional Text
2025 Β· Hoang-Bao Le, Allie Tran, Binh T. Nguyen, et al.
Abstract
Image-Guided Retrieval with Optional Text (IGROT) unifies visual retrieval (without text) and composed retrieval (with text). Despite its relevance in applications like Google Image and Bing, progress has been limited by the lack of an accessible benchmark and methods that balance performance across subtasks. Large-scale datasets such as MagicLens are comprehensive but computationally prohibitive, while existing models often favor either visual or compositional queries. We introduce FIGROTD, a lightweight yet high-quality IGROT dataset with 16,474 training triplets and 1,262 test triplets across CIR, SBIR, and CSTBIR. To reduce redundancy, we propose the Variance Guided Feature Mask (VaGFeM), which selectively enhances discriminative dimensions based on variance statistics. We further adopt a dual-loss design (InfoNCE + Triplet) to improve compositional reasoning. Trained on FIGROTD, VaGFeM achieves competitive results on nine benchmarks, reaching 34.8 mAP@10 on CIRCO and 75.7 mAP@200
Authors
(none)
Tags
Stats
Related papers
- UNION: A Lightweight Target Representation For Efficient Zero-shot Image-guided Retrieval With Optional Textual Queries (2025)0.00
- DVF: Advancing Robust And Accurate Fine-grained Image Retrieval With Retrieval Guidelines (2024)9.03
- Pseudo-triplet Guided Few-shot Composed Image Retrieval (2024)0.00
- Benchmark Granularity And Model Robustness For Image-text Retrieval (2024)0.00
- Fico-itr: Bridging Fine-grained And Coarse-grained Image-text Retrieval For Comparative Performance Analysis (2024)3.58
- Instance-level Composed Image Retrieval (2025)0.00
- Few Shots Text To Image Retrieval: New Benchmarking Dataset And Optimization Methods (2026)0.00
- Automatic Synthesis Of High-quality Triplet Data For Composed Image Retrieval (2025)0.00