← all papers · overview

Telerag: Efficient Retrieval-augmented Generation Inference With Lookahead Retrieval

Abstract

Retrieval-augmented generation (RAG) extends large language models (LLMs) with external data sources to enhance factual correctness and domain coverage. Modern RAG pipelines rely on large datastores, creating a significant system challenge: achieving high throughput and low latency is difficult, especially when GPU memory is limited. To address these challenges, we propose TeleRAG, an efficient in

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).