← all papers · overview

Satori: Reinforcement Learning With Chain-of-action-thought Enhances LLM Reasoning Via Autoregressive Search

Abstract

Large language models (LLMs) have demonstrated remarkable reasoning capabilities across diverse domains. Recent studies have shown that increasing test-time computation enhances LLMs' reasoning capabilities. This typically involves extensive sampling at inference time guided by an external LLM verifier, resulting in a two-player system. Despite external guidance, the effectiveness of this system d

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).