← all papers · overview

Understanding Long Videos With Multimodal Language Models

Abstract

Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs influence this strong performance. Surprisingly, we discover that LLM-based approaches can yield surprisingly good accuracy on long-video tasks with limited video in

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).