← all papers · overview

Hierarchical Adaptive Eviction For KV Cache Management In Multimodal Language Models

Abstract

The integration of visual information into Large Language Models (LLMs) has enabled Multimodal LLMs (MLLMs), but the quadratic memory and computational costs of Transformer architectures remain a bottleneck. Existing KV cache eviction strategies fail to address the heterogeneous attention distributions between visual and text tokens, leading to suboptimal efficiency or degraded performance. In thi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).