← all papers · overview

Zero Sum SVD: Balancing Loss Sensitivity For Low Rank LLM Compression

Abstract

Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD-based compression reduces storage and can speed up inference via low-rank factors, yet performance depends on how rank is allocated under a global compression ratio. Prior methods often use homogeneous ranks for similarly sized matrices, despite large

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).