← all papers · overview

Proxy-rlhf: Decoupling Generation And Alignment In Large Language Model With Proxy

Abstract

Reinforcement Learning from Human Feedback (RLHF) is the prevailing approach to ensure Large Language Models (LLMs) align with human values. However, existing RLHF methods require a high computational cost, one main reason being that RLHF assigns both the generation and alignment tasks to the LLM simultaneously. In this paper, we introduce Proxy-RLHF, which decouples the generation and alignment p

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).