← all papers · overview

RAPO: Expanding Exploration For LLM Agents Via Retrieval-augmented Policy Optimization

Abstract

Agentic Reinforcement Learning (Agentic RL) has shown remarkable potential in large language model-based (LLM) agents. These works can empower LLM agents to tackle complex tasks via multi-step, tool-integrated reasoning. However, an inherent limitation of existing Agentic RL methods is their reliance on a pure on-policy paradigm for exploration, restricting exploration to the agent's self-generate

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).