← all papers · overview

Thickening-to-thinning: Reward Shaping Via Human-inspired Learning Dynamics For LLM Reasoning

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for enhancing reasoning in Large Language Models (LLMs). However, it frequently encounters challenges such as entropy collapse, excessive verbosity, and insufficient exploration for hard problems. Crucially, existing reward schemes fail to distinguish between the need for extensive search during problem-solvi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).