← all papers · overview

Q Cache: Visual Attention Is Valuable In Less Than Half Of Decode Layers For Multimodal Large Language Model

Abstract

Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual tokens engenders a substantial computational load and key-value (KV) cache footprint bottleneck. Existing approaches focus on token-wise optimization, leveraging diverse intricate token pruning techniques to eliminate non-

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).