← all papers · overview

Gpt4video: A Unified Multimodal Large Language Model For Lnstruction-followed Understanding And Safety-aware Generation

Abstract

While the recent advances in Multimodal Large Language Models (MLLMs) constitute a significant leap forward in the field, these models are predominantly confined to the realm of input-side multimodal comprehension, lacking the capacity for multimodal content generation. To fill this gap, we present GPT4Video, a unified multi-model framework that empowers Large Language Models (LLMs) with the capab

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).