← all papers · overview

Vision-language Instruction Tuning: A Review And Analysis

Abstract

Instruction tuning is a crucial supervised training phase in Large Language Models (LLMs), aiming to enhance the LLM's ability to generalize instruction execution and adapt to user preferences. With the increasing integration of multi-modal data into LLMs, there is growing interest in Vision-Language Instruction Tuning (VLIT), which presents more complex characteristics compared to pure text instr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).