← all papers · overview

Flexllm: Token-level Co-serving Of LLM Inference And Finetuning With SLO Guarantees

Abstract

Finetuning large language models (LLMs) is essential for task adaptation, yet today's serving stacks isolate inference and finetuning on separate GPU clusters -- wasting resources and under-utilizing hardware. We introduce FlexLLM, the first system to co-serve LLM inference and PEFT-based finetuning on shared GPUs by fusing computation at the token level. FlexLLM's static compilation optimizations

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).