← all papers · overview

Adaptive Inference-time Compute: Llms Can Predict If They Can Do Better, Even Mid-generation

Abstract

Inference-time computation is a powerful paradigm to enhance the performance of large language models (LLMs), with Best-of-N sampling being a widely used technique. However, this method is computationally expensive, requiring both (1) an external reward model and (2) the generation of multiple samples. In this work, we introduce a new generative self-evaluation scheme designed to adaptively reduce

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).