← all papers · overview

Implicit Multimodal Alignment: On The Generalization Of Frozen Llms To Multimodal Inputs

Abstract

Large Language Models (LLMs) have demonstrated impressive performance on multimodal tasks, without any multimodal finetuning. They are the building block for Large Multimodal Models, yet, we still lack a proper understanding of their success. In this work, we expose frozen LLMs to image, video, audio and text inputs and analyse their internal representation aiming to understand their generalizatio

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).