← all papers · overview

Faster WIND: Accelerating Iterative Best-of- Distillation For LLM Alignment

Abstract

Recent advances in aligning large language models with human preferences have corroborated the growing importance of best-of-N distillation (BOND). However, the iterative BOND algorithm is prohibitively expensive in practice due to the sample and computation inefficiency. This paper addresses the problem by revealing a unified game-theoretic connection between iterative BOND and self-play alignmen

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).