← all papers · overview

Beyond Test-time Compute Strategies: Advocating Energy-per-token In LLM Inference

Abstract

Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks but come with substantial energy and computational costs, particularly in request-heavy scenarios. In many real-world applications, the full scale and capabilities of LLMs are often unnecessary, as Small Language Models (SLMs) can provide accurate responses for simpler text generation tasks. When enhanced with ad

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).