← all papers · overview

Elastic On-device LLM Service

Abstract

On-device Large Language Models (LLMs) are transforming mobile AI, catalyzing applications like UI automation without privacy concerns. Nowadays the common practice is to deploy a single yet powerful LLM as a general task solver for multiple requests. We identify a key system challenge in this paradigm: current LLMs lack the elasticity to serve requests that have diversified Service-Level Objectiv

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).