← all papers · overview

Ubench: Benchmarking Uncertainty In Large Language Models With Multiple Choice Questions

Abstract

Despite recent progress in systematic evaluation frameworks, benchmarking the uncertainty of large language models (LLMs) remains a highly challenging task. Existing methods for benchmarking the uncertainty of LLMs face three key challenges: the need for internal model access, additional training, or high computational costs. This is particularly unfavorable for closed-source models. To this end,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).