← all papers · overview

Fast Heterogeneous Serving: Scalable Mixed-scale LLM Allocation For Slo-constrained Inference

Abstract

Deploying large language model (LLM) inference at scale requires jointly selecting base models, provisioning heterogeneous GPUs, configuring parallelism, and distributing workloads under tight latency, accuracy, and budget constraints. Exact mixed-integer linear programming (MILP) approaches guarantee optimality but scale poorly. We propose two constraint-aware heuristics: a Greedy Heuristic (GH)

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).