← all papers · overview

Prompt Compression In The Wild: Measuring Latency, Rate Adherence, And Quality For Faster LLM Inference

Abstract

With the wide adoption of language models for IR -- and specifically RAG systems -- the latency of the underlying LLM becomes a crucial bottleneck, since the long contexts of retrieved passages lead large prompts and therefore, compute increase. Prompt compression, which reduces the size of input prompts while aiming to preserve performance on downstream tasks, has established itself as a cost-eff

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).