← all papers · overview

Enabling High-sparsity Foundational Llama Models With Efficient Pretraining And Deployment

Abstract

Large language models (LLMs) have revolutionized Natural Language Processing (NLP), but their size creates computational bottlenecks. We introduce a novel approach to create accurate, sparse foundational versions of performant LLMs that achieve full accuracy recovery for fine-tuning tasks at up to 70% sparsity. We achieve this for the LLaMA-2 7B model by combining the SparseGPT one-shot pruning me

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).