← all papers · overview

Wings: Learning Multimodal Llms Without Text-only Forgetting

Abstract

Multimodal large language models (MLLMs), initiated with a trained LLM, first align images with text and then fine-tune on multimodal mixed inputs. However, the MLLM catastrophically forgets the text-only instructions, which do not include images and can be addressed within the initial LLM. In this paper, we present Wings, a novel MLLM that excels in both text-only dialogues and multimodal compreh

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).