← all papers · overview

Vidcompress: Memory-enhanced Temporal Compression For Video Understanding In Large Language Models

Abstract

Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks. However, most Video-LLMs treat videos as a sequential set of individual frames, which results in insufficient temporal-spatial interaction that hinders fine-grained comprehension and difficulty in processing longer videos due to limited visual token capacity. To address these chal

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).