← all papers · overview

Better Instruction-following Through Minimum Bayes Risk

Abstract

General-purpose LLM judges capable of human-level evaluation provide not only a scalable and accurate way of evaluating instruction-following LLMs but also new avenues for supervising and improving their performance. One promising way of leveraging LLM judges for supervision is through Minimum Bayes Risk (MBR) decoding, which uses a reference-based evaluator to select a high-quality output from am

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).