← all papers · overview

SED-SFT: Selectively Encouraging Diversity In Supervised Fine-tuning

Abstract

Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has emerged as the standard post-training paradigm for large language models (LLMs). However, the conventional SFT process, driven by Cross-Entropy (CE) loss, often induces mode collapse, where models over-concentrate on specific response patterns. This lack of distributional diversity severely restricts the exploration efficienc

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).