← all papers · overview

Lococo: Dropping In Convolutions For Long Context Compression

Abstract

This paper tackles the memory hurdle of processing long context sequences in Large Language Models (LLMs), by presenting a novel approach, Dropping In Convolutions for Long Context Compression (LoCoCo). LoCoCo employs only a fixed-size Key-Value (KV) cache, and can enhance efficiency in both inference and fine-tuning stages. Diverging from prior methods that selectively drop KV pairs based on heur

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).