← all papers · overview

Laser: Learning To Adaptively Select Reward Models With Multi-armed Bandits

Abstract

Reward Models (RMs) are crucial to aligning large language models (LLMs), but the degree to which an RM specialized to one task (e.g. writing) generalizes to new tasks (e.g. math) is often not known a priori, often making using only one fixed RM to train LLMs suboptimal. However, optimizing LLMs with multiple RMs simultaneously can incur a prohibitively high computational cost and lead to conflict

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).