← all papers · overview

Sortedrl: Accelerating RL Training For Llms Through Online Length-aware Scheduling

Abstract

Scaling reinforcement learning (RL) has shown strong promise for enhancing the reasoning abilities of large language models (LLMs), particularly in tasks requiring long chain-of-thought generation. However, RL training efficiency is often bottlenecked by the rollout phase, which can account for up to 70% of total training time when generating long trajectories (e.g., 16k tokens), due to slow autor

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).