← all papers · overview

Arclight: A Lightweight LLM Inference Architecture For Many-core Cpus

Abstract

Although existing frameworks for large language model (LLM) inference on CPUs are mature, they fail to fully exploit the computation potential of many-core CPU platforms. Many-core CPUs are widely deployed in web servers and high-end networking devices, and are typically organized into multiple NUMA nodes that group cores and memory. Current frameworks largely overlook the substantial overhead of

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).