← all papers · overview

Intuitive Fine-tuning: Towards Simplifying Alignment Into A Single Process

Abstract

Supervised Fine-Tuning (SFT) and Preference Optimization (PO) are key processes for aligning Language Models (LMs) with human preferences post pre-training. While SFT excels in efficiency and PO in effectiveness, they are often combined sequentially without integrating their optimization objectives. This approach ignores the opportunities to bridge their paradigm gap and take the strengths from bo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).