← all papers · overview

Learning Adaptive LLM Decoding

Abstract

Decoding from large language models (LLMs) typically relies on fixed sampling hyperparameters (e.g., temperature, top-p), despite substantial variation in task difficulty and uncertainty across prompts and individual decoding steps. We propose to learn adaptive decoding policies that dynamically select sampling strategies at inference time, conditioned on available compute resources. Rather than f

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).