Imagebert: Cross-modal Pre-training With Large-scale Weak-supervised Image-text Data
2020 Β· di Qi, Lin Su, Jia Song, et al.
Abstract
In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different modalities as input and models the relationship between them. The model is pre-trained on four tasks simultaneously: Masked Language Modeling (MLM), Masked Object Classification (MOC), Masked Region Feature Regression (MRFR), and Image Text Matching (ITM). To further enhance the pre-training quality, we have collected a Large-scale weAk-supervised Image-Text (LAIT) dataset from Web. We first pre-train the model on this dataset, then conduct a second stage pre-training on Conceptual Captions and SBU Captions. Our experiments show that multi-stage pre-training strategy outperforms single-stage pre-training. We also fine-tune and evaluate our pre-trained ImageBERT model on image retrieval and text retrieval tasks, and achieve new state-of-the-art results on both MSCOCO and Flickr30k datasets.
Authors
(none)
Tags
Stats
Related papers
- Vilbert: Pretraining Task-agnostic Visiolinguistic Representations For Vision-and-language Tasks (2019)0.00
- Mllms-augmented Visual-language Representation Learning (2023)0.00
- UC2: Universal Cross-lingual Cross-modal Vision-and-language Pre-training (2021)13.05
- Lexlip: Lexicon-bottlenecked Language-image Pre-training For Large-scale Image-text Retrieval (2023)10.85
- Lightningdot: Pre-training Visual-semantic Embeddings For Real-time Image-text Retrieval (2021)17.42
- Dreamlip: Language-image Pre-training With Long Captions (2024)10.61
- COTS: Collaborative Two-stream Vision-language Pre-training Model For Cross-modal Retrieval (2022)13.60
- MILES: Visual BERT Pre-training With Injected Language Semantics For Video-text Retrieval (2022)10.61