← all papers · overview

Tangent Space Fine-tuning For Directional Preference Alignment In Large Language Models

Abstract

Our goal is to enable large language models (LLMs) to balance multiple human preference dimensions; such as helpfulness, safety, and verbosity, through principled and controllable alignment. Existing preference optimization methods, including Direct Preference Optimization (DPO), collapse feedback into a single scalar reward, fixing one balance among objectives and preventing traversal of the Pare

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).