← all papers · overview

Entropyprune: Matrix Entropy Guided Visual Token Pruning For Multimodal Large Language Models

Abstract

Multimodal large language models (MLLMs) incur substantial inference cost due to the processing of hundreds of visual tokens per image. Although token pruning has proven effective for accelerating inference, determining when and where to prune remains largely heuristic. Existing approaches typically rely on static, empirically selected layers, which limit interpretability and transferability acros

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).