← all papers · overview

SHIELD: A Segmented Hierarchical Memory Architecture For Energy-efficient LLM Inference On Edge Npus

Abstract

Large Language Model (LLM) inference on edge Neural Processing Units (NPUs) is fundamentally constrained by limited on-chip memory capacity. Although high-density embedded DRAM (eDRAM) is attractive for storing activation workspaces, its periodic refresh consumes substantial energy. Prior work has primarily focused on reducing off-chip traffic or optimizing refresh for persistent Key-Value (KV) ca

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).