← all papers · overview

Probe Pruning: Accelerating Llms Through Dynamic Pruning Via Model-probing

Abstract

We introduce Probe Pruning (PP), a novel framework for online, dynamic, structured pruning of Large Language Models (LLMs) applied in a batch-wise manner. PP leverages the insight that not all samples and tokens contribute equally to the model's output, and probing a small portion of each batch effectively identifies crucial weights, enabling tailored dynamic pruning for different batches. It comp

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).