← all papers · overview

LSAQ: Layer-specific Adaptive Quantization For Large Language Model Deployment

Abstract

As Large Language Models (LLMs) demonstrate exceptional performance across various domains, deploying LLMs on edge devices has emerged as a new trend. Quantization techniques, which reduce the size and memory requirements of LLMs, are effective for deploying LLMs on resource-limited edge devices. However, existing one-size-fits-all quantization methods often fail to dynamically adjust the memory r

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).