← all papers · overview

Value Augmented Sampling For Language Model Alignment And Personalization

Abstract

Aligning Large Language Models (LLMs) to cater to different human preferences, learning new skills, and unlearning harmful behavior is an important problem. Search-based methods, such as Best-of-N or Monte-Carlo Tree Search, are performant, but impractical for LLM adaptation due to their high inference cost. On the other hand, using Reinforcement Learning (RL) for adaptation is computationally eff

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).