← all papers · overview

SVD Contextual Sparsity Predictors For Fast LLM Inference

Abstract

Contextual sparsity is one of the approaches used to reduce computational complexity in the inference process of large language models (LLMs). Existing techniques for efficient LLM inference acceleration based on contextual sparsity with minimal accuracy degradation require training sparse pattern predictors. This paper presents a framework for accelerating inference of ReGLU-based feed-forward ne

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).