← all papers · overview

Spotserve: Serving Generative Large Language Models On Preemptible Instances

Abstract

The high computational and memory requirements of generative large language models (LLMs) make it challenging to serve them cheaply. This paper aims to reduce the monetary cost for serving LLMs by leveraging preemptible GPU instances on modern clouds, which offer accesses to spare GPUs at a much cheaper price than regular instances but may be preempted by the cloud at any time. Serving LLMs on pre

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).