← all papers · overview

Turbospec: Closed-loop Speculation Control System For Optimizing LLM Serving Goodput

Abstract

Large Language Model (LLM) serving systems batch concurrent user requests to achieve efficient serving. However, in real-world deployments, such inter-request parallelism from batching is often limited by external factors such as low request rates or memory constraints. Recent works focus on intra-request parallelism from speculative decoding as a solution to this problem. Unfortunately, benefits

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).