← all papers · overview

Llmcad: Fast And Scalable On-device Large Language Model Inference

Abstract

Generative tasks, such as text generation and question answering, hold a crucial position in the realm of mobile applications. Due to their sensitivity to privacy concerns, there is a growing demand for their execution directly on mobile devices. Currently, the execution of these generative tasks heavily depends on Large Language Models (LLMs). Nevertheless, the limited memory capacity of these de

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).