← all papers · overview

Effective Moe-based LLM Compression By Exploiting Heterogeneous Inter-group Experts Routing Frequency And Information Density

Abstract

Mixture-of-Experts (MoE) based Large Language Models (LLMs) have achieved superior performance, yet the massive memory overhead caused by storing multiple expert networks severely hinders their practical deployment. Singular Value Decomposition (SVD)-based compression has emerged as a promising post-training technique; however, most existing methods apply uniform rank allocation or rely solely on

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).