← all papers · overview

Crayon: Customized On-device LLM Via Instant Adapter Blending And Edge-server Hybrid Inference

Abstract

The customization of large language models (LLMs) for user-specified tasks gets important. However, maintaining all the customized LLMs on cloud servers incurs substantial memory and computational overheads, and uploading user data can also lead to privacy concerns. On-device LLMs can offer a promising solution by mitigating these issues. Yet, the performance of on-device LLMs is inherently constr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).