← all papers · overview

Towards Coarse-to-fine Evaluation Of Inference Efficiency For Large Language Models

Abstract

In real world, large language models (LLMs) can serve as the assistant to help users accomplish their jobs, and also support the development of advanced applications. For the wide application of LLMs, the inference efficiency is an essential concern, which has been widely studied in existing work, and numerous optimization algorithms and code libraries have been proposed to improve it. Nonetheless

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).