← all papers · overview

Optimus: Accelerating Large-scale Multi-modal LLM Training By Bubble Exploitation

Abstract

Multimodal large language models (MLLMs) have extended the success of large language models (LLMs) to multiple data types, such as image, text and audio, achieving significant performance in various domains, including multimodal translation, visual question answering and content generation. Nonetheless, existing systems are inefficient to train MLLMs due to substantial GPU bubbles caused by the he

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).