Learning To Embed Semantic Similarity For Joint Image-text Retrieval
2022 Β· Noam Malali, Yosi Keller
Abstract
We present a deep learning approach for learning the joint semantic embeddings of images and captions in a Euclidean space, such that the semantic similarity is approximated by the L2 distances in the embedding space. For that, we introduce a metric learning scheme that utilizes multitask learning to learn the embedding of identical semantic concepts using a center loss. By introducing a differentiable quantization scheme into the end-to-end trainable network, we derive a semantic embedding of semantically similar concepts in Euclidean space. We also propose a novel metric learning formulation using an adaptive margin hinge loss, that is refined during the training phase. The proposed scheme was applied to the MS-COCO, Flicke30K and Flickr8K datasets, and was shown to compare favorably with contemporary state-of-the-art approaches.
Authors
(none)
Tags
Stats
Related papers
- Learning To Learn From Web Data Through Deep Semantic Embeddings (2018)9.03
- Deep Multimodal Image-text Embeddings For Automatic Cross-media Retrieval (2020)0.00
- Webly Supervised Joint Embedding For Cross-modal Image-text Retrieval (2018)13.17
- Learning Robust Visual-semantic Embeddings (2017)15.22
- Hierarchy-based Image Embeddings For Semantic Image Retrieval (2018)13.84
- Image Search Using Multilingual Texts: A Cross-modal Learning Approach Between Image And Text (2019)0.00
- Aligning Multilingual Word Embeddings For Cross-modal Retrieval Task (2019)2.26
- Divide And Conquer The Embedding Space For Metric Learning (2019)14.39