← all papers · overview

Constraint-rectified Training For Efficient Chain-of-thought

Abstract

Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), especially when combined with reinforcement learning (RL) based post-training methods. While longer reasoning traces can improve answer quality and unlock abilities such as self-correction, they also incur high inference costs and often introduce redundant steps, known as overthinking. Rec

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).