← all papers · overview

Accelgen: Heterogeneous Slo-guaranteed High-throughput LLM Inference Serving For Diverse Applications

Abstract

In this paper, we consider a mixed-prompt scenario for a large language model (LLM) inference serving system that supports diverse applications with both short prompts and long prompts and heterogeneous SLOs for iteration time. To improve throughput when handling long prompts, previous research introduces a chunking method, but has not addressed heterogeneous SLOs. To address the limitation, we pr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).