← all papers · overview

Best-of-tails: Bridging Optimism And Pessimism In Inference-time Alignment

Abstract

Inference-time alignment effectively steers large language models (LLMs) by generating multiple candidates from a reference model and selecting among them with an imperfect reward model. However, current strategies face a fundamental dilemma: ``optimistic'' approaches like Best-of- suffer from reward hacking, while ``pessimistic'' regularized methods often stifle the exploration needed to dis

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).