← all papers · overview

Loma: Lossless Compressed Memory Attention

Abstract

Large Language Models (LLMs) face limitations due to the high demand on GPU memory and computational resources when handling long contexts. While sparsify the Key-Value (KV) cache of transformer model is a typical strategy to alleviate resource usage, it unavoidably results in the loss of information. We introduce Lossless Compressed Memory Attention (LoMA), a novel approach that enables lossless

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).