← all papers · overview

QCQA: Quality And Capacity-aware Grouped Query Attention

Abstract

Excessive memory requirements of key and value features (KV-cache) present significant challenges in the autoregressive inference of large language models (LLMs), restricting both the speed and length of text generation. Approaches such as Multi-Query Attention (MQA) and Grouped Query Attention (GQA) mitigate these challenges by grouping query heads and consequently reducing the number of correspo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).