← all papers · overview

Calibrating Long-form Generations From Large Language Models

Abstract

To enhance Large Language Models' (LLMs) reliability, calibration is essential -- the model's assessed confidence scores should align with the actual likelihood of its responses being correct. However, current confidence elicitation methods and calibration metrics typically rely on a binary true/false assessment of response correctness. This approach does not apply to long-form generation, where a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).