← all papers · overview

To 2:4 Sparsity And Beyond: Neuron-level Activation Function To Accelerate LLM Pre-training

Abstract

Trainings of Large Language Models are generally bottlenecked by matrix multiplications. In the Transformer architecture, a large portion of these operations happens in the Feed Forward Network (FFN), and this portion increases for larger models, up to 50% of the total pretraining floating point operations. We show that we can leverage hardware-accelerated sparsity to accelerate all matrix multipl

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).