← all papers · overview

Infmllm: A Unified Framework For Visual-language Tasks

Abstract

Large language models (LLMs) have proven their remarkable versatility in handling a comprehensive range of language-centric applications. To expand LLMs' capabilities to a broader spectrum of modal inputs, multimodal large language models (MLLMs) have attracted growing interest. This work delves into enabling LLMs to tackle more vision-language-related tasks, particularly image captioning, visual

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).