← all papers · overview

Meancache: User-centric Semantic Caching For LLM Web Services

Abstract

Large Language Models (LLMs) like ChatGPT and Llama have revolutionized natural language processing and search engine dynamics. However, these models incur exceptionally high computational costs. For instance, GPT-3 consists of 175 billion parameters, where inference demands billions of floating-point operations. Caching is a natural solution to reduce LLM inference costs on repeated queries, whic

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).