← all papers · overview

Prefillshare: A Shared Prefill Module For KV Reuse In Multi-llm Disaggregated Serving

Abstract

Multi-agent systems increasingly orchestrate multiple specialized language models to solve complex real-world problems, often invoking them over a shared context. This execution pattern repeatedly processes the same prompt prefix across models. Consequently, each model redundantly executes the prefill stage and maintains its own key-value (KV) cache, increasing aggregate prefill load and worsening

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).