← all papers · overview

Ragcache: Efficient Knowledge Caching For Retrieval-augmented Generation

Abstract

Retrieval-Augmented Generation (RAG) has shown significant improvements in various natural language processing tasks by integrating the strengths of large language models (LLMs) and external knowledge databases. However, RAG introduces long sequence generation and leads to high computation and memory costs. We propose RAGCache, a novel multilevel dynamic caching system tailored for RAG. Our analys

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).