← all papers · overview

Predicting And Improving Test-time Scaling Laws Via Reward Tail-guided Search

Abstract

Test-time scaling has emerged as a critical avenue for enhancing the reasoning capabilities of Large Language Models (LLMs). Though the straight-forward ''best-of-'' (BoN) strategy has already demonstrated significant improvements in performance, it lacks principled guidance on the choice of , budget allocation, and multi-stage decision-making, thereby leaving substantial room for optimi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).