← all papers · overview

GVPO: Group Variance Policy Optimization For Large Language Model Post-training

Abstract

Post-training plays a crucial role in refining and aligning large language models to meet specific tasks and human preferences. While recent advancements in post-training techniques, such as Group Relative Policy Optimization (GRPO), leverage increased sampling with relative reward scoring to achieve superior performance, these methods often suffer from training instability that limits their pract

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).