← all papers · overview

PAT: Pruning-aware Tuning For Large Language Models

Abstract

Large language models (LLMs) excel in language tasks, especially with supervised fine-tuning after pre-training. However, their substantial memory and computational requirements hinder practical applications. Structural pruning, which reduces less significant weight dimensions, is one solution. Yet, traditional post-hoc pruning often leads to significant performance loss, with limited recovery fro

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).