← all papers · overview

SCOPE: Saliency-coverage Oriented Token Pruning For Efficient Multimodel Llms

Abstract

Multimodal Large Language Models (MLLMs) typically process a large number of visual tokens, leading to considerable computational overhead, even though many of these tokens are redundant. Existing visual token pruning methods primarily focus on selecting the most salient tokens based on attention scores, resulting in the semantic incompleteness of the selected tokens. In this paper, we propose a novel visual token pruning strategy, called \textbf\{S\}aliency-\textbf\{C\}overage \textbf\{O\}riented token \textbf\{P\}runing for \textbf\{E\}fficient MLLMs (SCOPE), to jointly model both the saliency and coverage of the selected visual tokens to better preserve semantic completeness. Specifically, we introduce a set-coverage for a given set of selected tokens, computed based on the token relationships. We then define a token-coverage gain for each unselected token, quantifying how much additional coverage would be obtained by including it. By integrating the saliency score into the token-co

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).