← all papers · overview

Aligngpt: Multi-modal Large Language Models With Adaptive Alignment Capability

Abstract

Multimodal Large Language Models (MLLMs) are widely regarded as crucial in the exploration of Artificial General Intelligence (AGI). The core of MLLMs lies in their capability to achieve cross-modal alignment. To attain this goal, current MLLMs typically follow a two-phase training paradigm: the pre-training phase and the instruction-tuning phase. Despite their success, there are shortcomings in t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).