Support-set Bottlenecks For Video-text Representation Learning
2020 Β· Mandela Patrick, Po-Yao Huang, Yuki Asano, et al.
Abstract
The dominant paradigm for learning video-text representations -- noise contrastive learning -- increases the similarity of the representations of pairs of samples that are known to be related, such as text and video from the same sample, and pushes away the representations of all other pairs. We posit that this last behaviour is too strict, enforcing dissimilar representations even for samples that are semantically-related -- for example, visually similar videos or ones that share the same depicted action. In this paper, we propose a novel method that alleviates this by leveraging a generative model to naturally push these related samples together: each sample's caption must be reconstructed as a weighted combination of other support samples' visual representations. This simple idea ensures that representations are not overly-specialized to individual samples, are reusable across the dataset, and results in representations that explicitly encode semantics shared between samples, unlike
Authors
(none)
Tags
Stats
Related papers
- Rebalancing Contrastive Alignment With Bottlenecked Semantic Increments In Text-video Retrieval (2025)1.69
- Self-supervised Video Representation Learning With Cross-stream Prototypical Contrasting (2021)8.82
- Expectation-maximization Contrastive Learning For Compact Video-and-language Representations (2022)2.26
- Contrastive Video-language Learning With Fine-grained Frame Sampling (2022)6.77
- Normalized Contrastive Learning For Text-video Retrieval (2022)6.77
- TCLR: Temporal Contrastive Learning For Video Representation (2021)15.78
- Lat: Latent Translation With Cycle-consistency For Video-text Retrieval (2022)0.00
- Unifying Latent And Lexicon Representations For Effective Video-text Retrieval (2024)0.00