← all papers · overview

Dart-ing Through The Drift: Dynamic Tracing Of Knowledge Neurons For Adaptive Inference-time Pruning

Abstract

Large Language Models (LLMs) exhibit substantial parameter redundancy, particularly in Feed-Forward Networks (FFNs). Existing pruning methods suffer from two primary limitations. First, reliance on dataset-specific calibration introduces significant data dependency and computational overhead. Second, being predominantly static, they fail to account for the evolving subset of knowledge neurons in L

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).