← all papers · overview

Beyond The Speculative Game: A Survey Of Speculative Execution In Large Language Models

Abstract

With the increasingly giant scales of (causal) large language models (LLMs), the inference efficiency comes as one of the core concerns along the improved performance. In contrast to the memory footprint, the latency bottleneck seems to be of greater importance as there can be billions of requests to a LLM (e.g., GPT-4) per day. The bottleneck is mainly due to the autoregressive innateness of LLMs

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).