← all papers · overview

Nli:non-uniform Linear Interpolation Approximation Of Nonlinear Operations For Efficient Llms Inference

Abstract

Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of tasks, but their deployment is often constrained by substantial memory footprints and computational costs. While prior work has achieved significant progress in compressing and accelerating linear layers, nonlinear layers-such as SiLU, RMSNorm, and Softmax-still heavily depend on high-precision floating-po

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).