← all papers · overview

Hygen: Efficient LLM Serving Via Elastic Online-offline Request Co-location

Abstract

Large language models (LLMs) have facilitated a wide range of applications with distinct service-level objectives (SLOs), from latency-sensitive online tasks like interactive chatbots to throughput-oriented offline workloads like data synthesis. The existing deployment model, which dedicates machines to each workload, simplifies SLO management but often leads to poor resource utilization. This pap

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).