← all papers · overview

Serverlessllm: Low-latency Serverless Inference For Large Language Models

Abstract

This paper presents ServerlessLLM, a distributed system designed to support low-latency serverless inference for Large Language Models (LLMs). By harnessing the substantial near-GPU storage and memory capacities of inference servers, ServerlessLLM achieves effective local checkpoint storage, minimizing the need for remote checkpoint downloads and ensuring efficient checkpoint loading. The design o

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).