← all papers · overview

From Image To Video, What Do We Need In Multimodal Llms?

Abstract

Covering from Image LLMs to the more complex Video LLMs, the Multimodal Large Language Models (MLLMs) have demonstrated profound capabilities in comprehending cross-modal information as numerous studies have illustrated. Previous methods delve into designing comprehensive Video LLMs through integrating video foundation models with primitive LLMs. Despite its effectiveness, such paradigm renders Vi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).