← all papers · overview

Detecting Overflow In Compressed Token Representations For Retrieval-augmented Generation

Abstract

Efficient long-context processing remains a crucial challenge for contemporary large language models (LLMs), especially in resource-constrained environments. Soft compression architectures promise to extend effective context length by replacing long token sequences with smaller sets of learned compressed tokens. Yet, the limits of compressibility -- and when compression begins to erase task-releva

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).