← all papers · overview

Deepstack: Deeply Stacking Visual Tokens Is Surprisingly Simple And Effective For Lmms

Abstract

Most large multimodal models (LMMs) are implemented by feeding visual tokens as a sequence into the first layer of a large language model (LLM). The resulting architecture is simple but significantly increases computation and memory costs, as it has to handle a large number of additional tokens in its input layer. This paper presents a new architecture DeepStack for LMMs. Considering layers

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).