← all papers · overview

Slo-aware Compute Resource Allocation For Prefill-decode Disaggregated LLM Inference

Abstract

Prefill-Decode (P/D) disaggregation has emerged as a widely adopted optimization strategy for Large Language Model (LLM) inference. However, there currently exists no well-established methodology for determining the optimal number of P/D hardware resources, subject to constraints on total throughput, service level objectives (SLOs), and request characteristics - specifically input and output lengt

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).