Cross-modal Search Method Of Technology Video Based On Adversarial Learning And Feature Fusion
2022 Β· Xiangbin Liu, Junping Du, Meiyu Liang, et al.
Abstract
Technology videos contain rich multi-modal information. In cross-modal information search, the data features of different modalities cannot be compared directly, so the semantic gap between different modalities is a key problem that needs to be solved. To address the above problems, this paper proposes a novel Feature Fusion based Adversarial Cross-modal Retrieval method (FFACR) to achieve text-to-video matching, ranking and searching. The proposed method uses the framework of adversarial learning to construct a video multimodal feature fusion network and a feature mapping network as generator, a modality discrimination network as discriminator. Multi-modal features of videos are obtained by the feature fusion network. The feature mapping network projects multi-modal features into the same semantic space based on semantics and similarity. The modality discrimination network is responsible for determining the original modality of features. Generator and discriminator are trained alterna
Authors
(none)
Tags
Stats
Related papers
- Everything At Once -- Multi-modal Fusion Transformer For Video Retrieval (2021)15.78
- Video And Audio Are Images: A Cross-modal Mixer For Original Data On Video-audio Retrieval (2023)7.16
- MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion (2025)2.26
- Towards Fast Adaptation Of Pretrained Contrastive Models For Multi-channel Video-language Retrieval (2022)7.50
- Joint Fusion And Encoding: Advancing Multimodal Retrieval From The Ground Up (2025)0.00
- Integrating Information Theory And Adversarial Learning For Cross-modal Retrieval (2021)10.97
- Dual Encoding For Video Retrieval By Text (2020)16.05
- A Feature-space Multimodal Data Augmentation Technique For Text-video Retrieval (2022)12.43