← all papers · overview

Gradually Compacting Large Language Models For Reasoning Like A Boiling Frog

Abstract

Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, but their substantial size often demands significant computational resources. To reduce resource consumption and accelerate inference, it is essential to eliminate redundant parameters without compromising performance. However, conventional pruning methods that directly remove such parameters often lead to a dramatic

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).