← all papers · overview

Edkm: An Efficient And Accurate Train-time Weight Clustering For Large Language Models

Abstract

Since Large Language Models or LLMs have demonstrated high-quality performance on many complex language tasks, there is a great interest in bringing these LLMs to mobile devices for faster responses and better privacy protection. However, the size of LLMs (i.e., billions of parameters) requires highly effective compression to fit into storage-limited devices. Among many compression techniques, wei

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).