← all papers · overview

Policy Of Thoughts: Scaling LLM Reasoning Via Test-time Policy Evolution

Abstract

Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption. Current test-time scaling methods treat execution feedback merely as an external signal for filtering or rewriting trajectories, without internalizing it to improve the underlying reasoning strategy. Inspired by Popper's epistemology of "conjectures and refutation

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).