← all papers · overview

Data-efficient Alignment Of Large Language Models With Human Feedback Through Natural Language

Abstract

Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF) leverages human preference signals that are in the form of ranking of response pairs to perform this alignment. However, human preference on LLM outputs can come in much richer forms including natural language, which ma

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).