LongVideoBench
Emerging15papers using it
2025first seen
Dataset Card for LongVideoBench Large multimodal models (LMMs) are handling increasingly longer and more complex inputs. However, few public benchmarks are available to assess these advancements. To address this, we introduce LongVideoBench, a question-answering benchmark with video-language interleaved inputs up to an
Papers using LongVideoBench (15)
- Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video UnderstandingKimi-VL Technical ReportProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process RewardWhere to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video UnderstandingEvent-Anchored Frame Selection for Effective Long-Video UnderstandingHiMu: Hierarchical Multimodal Frame Selection for Long Video Question AnsweringReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video UnderstandingMSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video UnderstandingLinMU: Multimodal Understanding Made LinearThink-Clip-Sample: Slow-Fast Frame Selection for Video UnderstandingLiViBench: An Omnimodal Benchmark for Interactive Livestream Video UnderstandingVSI: Visual Subtitle Integration For Keyframe Selection To Enhance Long Video UnderstandingLess Is More: Token-efficient Video-qa Via Adaptive Frame-pruning And Semantic Graph IntegrationQ-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMsExploring the Effect of Reinforcement Learning on Video Understanding:
Insights from SEED-Bench-R1