← all papers · overview

Hire: High Recall Approximate Top- Estimation For Efficient LLM Inference

Abstract

Autoregressive decoding with generative Large Language Models (LLMs) on accelerators (GPUs/TPUs) is often memory-bound where most of the time is spent on transferring model parameters from high bandwidth memory (HBM) to cache. On the other hand, recent works show that LLMs can maintain quality with significant sparsity/redundancy in the feedforward (FFN) layers by appropriately training the model

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).