← all papers · overview

Preble: Efficient Distributed Prompt Scheduling For LLM Serving

Abstract

Prompts to large language models (LLMs) have evolved beyond simple user questions. For LLMs to solve complex problems, today's practices are to include domain-specific instructions, illustration of tool usages, and/or long context such as textbook chapters in prompts. As such, many parts of prompts are repetitive across requests. Recent works propose to cache and reuse KV state of prompts. However

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).