← all papers · overview

Asynchronous Verified Semantic Caching For Tiered LLM Architectures

Abstract

Large language models (LLMs) now sit in the critical path of search, assistance, and agentic workflows, making semantic caching essential for reducing inference cost and latency. Production deployments typically use a tiered static-dynamic design: a static cache of curated, offline vetted responses mined from logs, backed by a dynamic cache populated online. In practice, both tiers are commonly go

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).