← all papers · overview

Truthfulness Despite Weak Supervision: Evaluating And Training Llms Using Peer Prediction

Abstract

The evaluation and post-training of large language models (LLMs) rely on supervision, but strong supervision for difficult tasks is often unavailable, especially when evaluating frontier models. In such cases, models are demonstrated to exploit evaluations built on such imperfect supervision, leading to deceptive results. However, underutilized in LLM research, a wealth of mechanism design researc

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).