← all papers · overview

Mélange: Cost Efficient Large Language Model Serving By Exploiting GPU Heterogeneity

Abstract

Large language models (LLMs) are increasingly integrated into many online services, yet they remain cost-prohibitive to deploy due to the requirement of expensive GPU instances. Prior work has addressed the high cost of LLM serving by improving the inference engine, but less attention has been given to selecting the most cost-efficient GPU type(s) for a specific LLM service. There is a large and g

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).