3D-CSL: Self-supervised 3D Context Similarity Learning For Near-duplicate Video Retrieval
2022 Β· Rui Deng, Qian Wu, Yuke Li
Abstract
In this paper, we introduce 3D-CSL, a compact pipeline for Near-Duplicate Video Retrieval (NDVR), and explore a novel self-supervised learning strategy for video similarity learning. Most previous methods only extract video spatial features from frames separately and then design kinds of complex mechanisms to learn the temporal correlations among frame features. However, parts of spatiotemporal dependencies have already been lost. To address this, our 3D-CSL extracts global spatiotemporal dependencies in videos end-to-end with a 3D transformer and find a good balance between efficiency and effectiveness by matching on clip-level. Furthermore, we propose a two-stage self-supervised similarity learning strategy to optimize the entire network. Firstly, we propose PredMAE to pretrain the 3D transformer with video prediction task; Secondly, ShotMix, a novel video-specific augmentation, and FCS loss, a novel triplet loss, are proposed further promote the similarity learning results. The expe
Authors
(none)
Tags
Stats
Related papers
- CNN Retrieval Based Unsupervised Metric Learning For Near-duplicated Video Retrieval (2021)0.00
- Self-supervised Video Similarity Learning (2023)13.04
- Semi-supervised 3D Video Information Retrieval With Deep Neural Network And Bi-directional Dynamic-time Warping Algorithm (2023)0.00
- Visil: Fine-grained Spatio-temporal Video Similarity Learning (2019)13.70
- Cycle-contrast For Self-supervised Video Representation Learning (2020)0.00
- Nearest-neighbor Inter-intra Contrastive Learning From Unlabeled Videos (2023)0.00
- Audio-based Near-duplicate Video Retrieval With Audio Similarity Learning (2020)7.16
- Self-supervised Video Retrieval Transformer Network (2021)0.00