← all papers · overview

Speculative Prefill: Turbocharging TTFT With Lightweight And Training-free Token Importance Estimation

Abstract

Improving time-to-first-token (TTFT) is an essentially important objective in modern large language model (LLM) inference engines. Optimizing TTFT directly results in higher maximal QPS and meets the requirements of many critical applications. However, boosting TTFT is notoriously challenging since it is compute-bounded and the performance bottleneck shifts from the self-attention that many prior

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).