← all papers · overview

Loongserve: Efficiently Serving Long-context Large Language Models With Elastic Sequence Parallelism

Abstract

The context window of large language models (LLMs) is rapidly increasing, leading to a huge variance in resource usage between different requests as well as between different phases of the same request. Restricted by static parallelism strategies, existing LLM serving systems cannot efficiently utilize the underlying resources to serve variable-length requests in different phases. To address this

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).