← all papers · overview

RC-GRPO: Reward-conditioned Group Relative Policy Optimization For Multi-turn Tool Calling Agents

Abstract

Multi-turn tool calling is challenging for Large Language Models (LLMs) because rewards are sparse and exploration is expensive. A common recipe, SFT followed by GRPO, can stall when within-group reward variation is low (e.g., more rollouts in a group receive the all 0 or all 1 reward), making the group-normalized advantage uninformative and yielding vanishing updates. To address this problem, we

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).