← all papers · overview

Kvshare: An LLM Service System With Efficient And Effective Multi-tenant KV Cache Reuse

Abstract

Recent advances in long-text understanding have pushed the context length of large language models (LLMs) up to one million tokens. It boosts LLMs's accuracy and reasoning capacity but causes exorbitant computational costs and unsatisfactory Time to First Token (TTFT). KV cache reuse, which reuses the exact same KV cache of prefixes and templates or shares similar ones but with extra selective rec

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).