Paired Cross-modal Data Augmentation For Fine-grained Image-to-text Retrieval
2022 Β· Hao Wang, Guosheng Lin, Steven C. H. Hoi, et al.
Abstract
This paper investigates an open research problem of generating text-image pairs to improve the training of fine-grained image-to-text cross-modal retrieval task, and proposes a novel framework for paired data augmentation by uncovering the hidden semantic information of StyleGAN2 model. Specifically, we first train a StyleGAN2 model on the given dataset. We then project the real images back to the latent space of StyleGAN2 to obtain the latent codes. To make the generated images manipulatable, we further introduce a latent space alignment module to learn the alignment between StyleGAN2 latent codes and the corresponding textual caption features. When we do online paired data augmentation, we first generate augmented text through random token replacement, then pass the augmented text into the latent space alignment module to output the latent codes, which are finally fed to StyleGAN2 to generate the augmented images. We evaluate the efficacy of our augmented data approach on two public
Authors
(none)
Tags
Stats
Related papers
- A Feature-space Multimodal Data Augmentation Technique For Text-video Retrieval (2022)12.43
- Mixgen: A New Multi-modal Data Augmentation (2022)14.47
- Look, Imagine And Match: Improving Textual-visual Cross-modal Retrieval With Generative Models (2017)18.52
- Webly Supervised Joint Embedding For Cross-modal Image-text Retrieval (2018)13.17
- Cross-modal RAG: Sub-dimensional Text-to-image Retrieval-augmented Generation (2025)0.00
- Text-to-image Generation Via Implicit Visual Guidance And Hypernetwork (2022)0.00
- Cross-modal Attribute Insertions For Assessing The Robustness Of Vision-and-language Learning (2023)2.00
- Style-aware Contrastive Learning For Multi-style Image Captioning (2023)5.84