← all papers · overview

High-throughput LLM Inference On Heterogeneous Clusters

Abstract

Nowadays, many companies possess various types of AI accelerators, forming heterogeneous clusters. Efficiently leveraging these clusters for high-throughput large language model (LLM) inference services can significantly reduce costs and expedite task processing. However, LLM inference on heterogeneous clusters presents two main challenges. Firstly, different deployment configurations can result i

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).