MVBench
Emerging13papers using it
2025first seen
MVBench is a benchmark dataset used to evaluate the performance of vision-language models (VLMs) in tasks related to video understanding and temporal reasoning, specifically focusing on counting repetitions in video clips.
Papers using MVBench (13)
- Pushupbench: Your VLM Is Not Good At Counting PushupsVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextLongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference OptimizationNot All Modalities Are Equal: Instruction-Aware Gating for Multimodal VideosClue Matters: Leveraging Latent Visual Clues to Empower Video ReasoningMACD: Model-Aware Contrastive Decoding via Counterfactual DataMSVBench: Towards Human-Level Evaluation of Multi-Shot Video GenerationImproving Video Question Answering through query-based frame selectionVideo Evidence to Reasoning Efficient Video Understanding via Explicit Evidence GroundingEnhancing Temporal Understanding In Video-llms Through Stacked Temporal Attention In Vision EncodersRo-bench: Large-scale Robustness Evaluation Of Mllms With Text-driven Counterfactual VideosSeeing Is Not Reasoning: Mvpbench For Graph-based Evaluation Of Multi-path Visual Physical CotGam-agent: Game-theoretic And Uncertainty-aware Collaboration For Complex Visual Reasoning