← all papers · overview

Relaycaching: Accelerating LLM Collaboration Via Decoding KV Cache Reuse

Abstract

The increasing complexity of AI tasks has shifted the paradigm from monolithic models toward multi-agent large language model (LLM) systems. However, these collaborative architectures introduce a critical bottleneck: redundant prefill computation for shared content generated by previous agents, which significantly increases KV cache memory usage and time-to-first-token (TTFT). While various KV cac

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).