← all papers · overview

Adaptive Decoding Via Test-time Policy Learning For Self-improving Generation

Abstract

Decoding strategies largely determine the quality of Large Language Model (LLM) outputs, yet widely used heuristics such as greedy or fixed temperature/top-p decoding are static and often task-agnostic, leading to suboptimal or inconsistent generation quality across domains that demand stylistic or structural flexibility. We introduce a reinforcement learning-based decoder sampler that treats deco

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).