← all papers · overview

Multi-turn Reinforcement Learning From Preference Human Feedback

Abstract

Reinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate remarkable abilities in various tasks. Existing methods work by emulating the preferences at the single decision (turn) level, limiting their capabilities in settings that require planning or multi-turn interactions to ach

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).