← all papers · overview

Rethinking Data Selection For Supervised Fine-tuning

Ming Shen·2024

Abstract

Although supervised finetuning (SFT) has emerged as an essential technique to align large language models with humans, it is considered superficial, with style learning being its nature. At the same time, recent works indicate the importance of data selection for SFT, showing that finetuning with high-quality and diverse subsets of the original dataset leads to superior downstream performance. In

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).