← all papers · overview

How To Bridge The Gap Between Modalities: Survey On Multimodal Large Language Model

Abstract

We explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more. MLLMs demonstrate capabilities such as generating image captions and answering image-based questions, bridging the gap towards real-world human-computer interactions and hinting at a potential pathway to artificial general intelligence. However, M

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).