← all papers · overview

Forkkv: Scaling Multi-lora Agent Serving Via Copy-on-write Disaggregated KV Cache

Abstract

The serving paradigm of large language models (LLMs) is rapidly shifting towards complex multi-agent workflows where specialized agents collaborate over massive shared contexts. While Low-Rank Adaptation (LoRA) enables the efficient co-hosting of these specialized agents on a single base model, it introduces a critical memory footprint bottleneck during serving. Specifically, unique LoRA activatio

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).