← all papers · overview

Temperature As A Meta-policy: Adaptive Temperature In LLM Reinforcement Learning

Abstract

Temperature is a crucial hyperparameter in large language models (LLMs), controlling the trade-off between exploration and exploitation during text generation. High temperatures encourage diverse but noisy outputs, while low temperatures produce focused outputs but may cause premature convergence. Yet static or heuristic temperature schedules fail to adapt to the dynamic demands of reinforcement l

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).