← all papers · overview

Self-selected Attention Span For Accelerating Large Language Model Inference

Abstract

Large language models (LLMs) can solve challenging tasks. However, their inference computation on modern GPUs is highly inefficient due to the increasing number of tokens they must attend to as they generate new ones. To address this inefficiency, we capitalize on LLMs' problem-solving capabilities to optimize their own inference-time efficiency. We demonstrate with two specific tasks: (a) evaluat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).