← all papers · overview

Preserving Diversity In Supervised Fine-tuning Of Large Language Models

Abstract

Large Language Models (LLMs) typically rely on Supervised Fine-Tuning (SFT) to specialize in downstream tasks, with the Cross Entropy (CE) loss being the de facto choice. However, CE maximizes the likelihood of observed data without accounting for alternative possibilities. As such, CE usually leads to reduced diversity in the model's outputs, which hinders further development that requires sampli

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).