← all papers · overview

Radial Networks: Dynamic Layer Routing For High-performance Large Language Models

Abstract

Large language models (LLMs) often struggle with strict memory, latency, and power demands. To meet these demands, various forms of dynamic sparsity have been proposed that reduce compute on an input-by-input basis. These methods improve over static methods by exploiting the variance across individual inputs, which has steadily grown with the exponential increase in training data. Yet, the increas

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).