← all papers · overview

Prune As You Generate: Online Rollout Pruning For Faster And Better RLVR

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs). However, methods such as GRPO and DAPO suffer from substantial computational cost, since they rely on sampling many rollouts for each prompt. Moreover, in RLVR the relative advantage is often sparse: many samples become nearly all-correct or all-incorrect, yi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).