← all papers · overview

ALISE: Accelerating Large Language Model Serving With Speculative Scheduling

Abstract

Large Language Models (LLMs) represent a revolutionary advancement in the contemporary landscape of artificial general intelligence (AGI). As exemplified by ChatGPT, LLM-based applications necessitate minimal response latency and maximal throughput for inference serving. However, due to the unpredictability of LLM execution, the first-come-first-serve (FCFS) scheduling policy employed by current L

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).