← all papers · overview

Cross-modal Projection In Multimodal Llms Doesn't Really Project Visual Attributes To Textual Space

Abstract

Multimodal large language models (MLLMs) like LLaVA and GPT-4(V) enable general-purpose conversations about images with the language modality. As off-the-shelf MLLMs may have limited capabilities on images from domains like dermatology and agriculture, they must be fine-tuned to unlock domain-specific applications. The prevalent architecture of current open-source MLLMs comprises two major modules

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).