← all papers · overview

Dialcot Meets PPO: Decomposing And Exploring Reasoning Paths In Smaller Language Models

Abstract

Chain-of-Thought (CoT) prompting has proven to be effective in enhancing the reasoning capabilities of Large Language Models (LLMs) with at least 100 billion parameters. However, it is ineffective or even detrimental when applied to reasoning tasks in Smaller Language Models (SLMs) with less than 10 billion parameters. To address this limitation, we introduce Dialogue-guided Chain-of-Thought (Dial

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).