← all papers · overview

Soft Sequence Policy Optimization

Abstract

A significant portion of recent research on Large Language Model (LLM) alignment focuses on developing new policy optimization methods based on Group Relative Policy Optimization (GRPO). Two prominent directions have emerged: (i) a shift toward sequence-level importance sampling weights that better align with the sequence-level rewards used in many tasks, and (ii) alternatives to PPO-style clippin

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).