← all papers · overview

Self-evolution Fine-tuning For Policy Optimization

Abstract

The alignment of large language models (LLMs) is crucial not only for unlocking their potential in specific tasks but also for ensuring that responses meet human expectations and adhere to safety and ethical principles. Current alignment methodologies face considerable challenges. For instance, supervised fine-tuning (SFT) requires extensive, high-quality annotated samples, while reinforcement lea

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).