← all papers · overview

Sparse Fine-tuning For Inference Acceleration Of Large Language Models

Abstract

We consider the problem of accurate sparse fine-tuning of large language models (LLMs), that is, fine-tuning pretrained LLMs on specialized tasks, while inducing sparsity in their weights. On the accuracy side, we observe that standard loss-based fine-tuning may fail to recover accuracy, especially at high sparsities. To address this, we perform a detailed study of distillation-type losses, determ

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).