← all papers · overview

SCALM: Towards Semantic Caching For Automated Chat Services With Large Language Models

Abstract

Large Language Models (LLMs) have become increasingly popular, transforming a wide range of applications across various domains. However, the real-world effectiveness of their query cache systems has not been thoroughly investigated. In this work, we for the first time conducted an analysis on real-world human-to-LLM interaction data, identifying key challenges in existing caching solutions for LL

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).