← all papers · overview

Accelerating Inference In Large Language Models With A Unified Layer Skipping Strategy

Abstract

Recently, dynamic computation methods have shown notable acceleration for Large Language Models (LLMs) by skipping several layers of computations through elaborate heuristics or additional predictors. However, in the decoding process of existing approaches, different samples are assigned different computational budgets, which cannot guarantee a stable and precise acceleration effect. Furthermore,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).