← all papers · overview

Toward Optimal LLM Alignments Using Two-player Games

Abstract

The standard Reinforcement Learning from Human Feedback (RLHF) framework primarily focuses on optimizing the performance of large language models using pre-collected prompts. However, collecting prompts that provide comprehensive coverage is both tedious and challenging, and often fails to include scenarios that LLMs need to improve on the most. In this paper, we investigate alignment through the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).