← all papers · overview

When Linear Attention Meets Autoregressive Decoding: Towards More Effective And Efficient Linearized Large Language Models

Abstract

Autoregressive Large Language Models (LLMs) have achieved impressive performance in language tasks but face two significant bottlenecks: (1) quadratic complexity in the attention module as the number of tokens increases, and (2) limited efficiency due to the sequential processing nature of autoregressive LLMs during generation. While linear attention and speculative decoding offer potential soluti

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).