← all papers · overview

Delta-come: Training-free Delta-compression With Mixed-precision For Large Language Models

Abstract

Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresponding delta weights, which are then compressed using low-rank or low-bit approaches to reduce costs

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).