← all papers · overview

Parameter-efficient Tuning Helps Language Model Alignment

Abstract

Aligning large language models (LLMs) with human preferences is essential for safe and useful LLMs. Previous works mainly adopt reinforcement learning (RLHF) and direct preference optimization (DPO) with human feedback for alignment. Nevertheless, they have certain drawbacks. One such limitation is that they can only align models with one preference at the training time (e.g., they cannot learn to

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).