← all papers · overview

BRIEF: Bridging Retrieval And Inference For Multi-hop Reasoning Via Compression

Abstract

Retrieval-augmented generation (RAG) can supplement large language models (LLMs) by integrating external knowledge. However, as the number of retrieved documents increases, the input length to LLMs grows linearly, causing a dramatic increase in latency and a degradation in long-context understanding. This is particularly serious for multi-hop questions that require a chain of reasoning across docu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).