← all papers · overview

Conv-basis: A New Paradigm For Efficient Attention Inference And Gradient Computation In Transformers

Abstract

The self-attention mechanism is the key to the success of transformers in recent Large Language Models (LLMs). However, the quadratic computational cost in the input sequence length is a notorious obstacle for further improvement and scalability in longer contexts. In this work, we leverage the convolution-like structure of attention matrices to develop an efficient approximation

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).