← all papers · overview

Tender: Accelerating Large Language Models Via Tensor Decomposition And Runtime Requantization

Abstract

Large language models (LLMs) demonstrate outstanding performance in various tasks in machine learning and have thus become one of the most important workloads in today's computing landscape. However, deploying LLM inference poses challenges due to the high compute and memory requirements stemming from the enormous model size and the difficulty of running it in the integer pipelines. In this paper,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).