← all papers · overview

Accelerating Large Language Models Through Partially Linear Feed-forward Network

Abstract

Large language models (LLMs) demonstrate remarkable capabilities but face deployment challenges due to their massive parameter counts. While existing compression techniques like pruning can reduce model size, it leads to significant accuracy degradation under high compression ratios. We present a novel perspective inspired by constant folding in compiler optimization. Our approach enables paramete

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).