← all papers · overview

Igrpo: Self-feedback-driven LLM Reasoning

Abstract

Large Language Models (LLMs) have shown promise in solving complex mathematical problems, yet they still fall short of producing accurate and consistent solutions. Reinforcement Learning (RL) is a framework for aligning these models with task-specific rewards, improving overall quality and reliability. Group Relative Policy Optimization (GRPO) is an efficient, value-function-free alternative to Pr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).