← all papers · overview

Meki: Memory-based Expert Knowledge Injection For Efficient LLM Scaling

Abstract

Scaling Large Language Models (LLMs) typically relies on increasing the number of parameters or test-time computations to boost performance. However, these strategies are impractical for edge device deployment due to limited RAM and NPU resources. Despite hardware constraints, deploying performant LLM on edge devices such as smartphone remains crucial for user experience. To address this, we propo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).