Abstract
Test-time scaling has emerged as a critical avenue for enhancing the reasoning capabilities of Large Language Models (LLMs). Though the straight-forward ''best-of-'' (BoN) strategy has already demonstrated significant improvements in performance, it lacks principled guidance on the choice of , budget allocation, and multi-stage decision-making, thereby leaving substantial room for optimi