Viseret: A Simple Yet Effective Approach To Moment Retrieval Via Fine-grained Video Segmentation
2021 Β· Aiden Seungjoon Lee, Hanseok Oh, Minjoon Seo
Abstract
Video-text retrieval has many real-world applications such as media analytics, surveillance, and robotics. This paper presents the 1st place solution to the video retrieval track of the ICCV VALUE Challenge 2021. We present a simple yet effective approach to jointly tackle two video-text retrieval tasks (video retrieval and video corpus moment retrieval) by leveraging the model trained only on the video retrieval task. In addition, we create an ensemble model that achieves the new state-of-the-art performance on all four datasets (TVr, How2r, YouCook2r, and VATEXr) presented in the VALUE Challenge.
Authors
(none)
Tags
Stats
Related papers
- Towards Efficient And Robust Moment Retrieval System: A Unified Framework For Multi-granularity Models And Temporal Reranking (2025)2.26
- Hybrid-learning Video Moment Retrieval Across Multi-domain Labels (2024)0.00
- Frame-wise Cross-modal Matching For Video Moment Retrieval (2020)13.17
- Improving Video Corpus Moment Retrieval With Partial Relevance Enhancement (2024)7.89
- Verve: Versatile Retrieval For Videos Via Unified Embeddings (2026)0.00
- Video Moment Retrieval With Text Query Considering Many-to-many Correspondence Using Potentially Relevant Pair (2021)0.00
- CLIP2TV: Align, Match And Distill For Video-text Retrieval (2021)0.00
- A Lightweight Moment Retrieval System With Global Re-ranking And Robust Adaptive Bidirectional Temporal Search (2025)3.58