← all papers · overview

Language Models Can Evaluate Themselves Via Probability Discrepancy

Abstract

In this paper, we initiate our discussion by demonstrating how Large Language Models (LLMs), when tasked with responding to queries, display a more even probability distribution in their answers if they are more adept, as opposed to their less skilled counterparts. Expanding on this foundational insight, we propose a new self-evaluation method ProbDiff for assessing the efficacy of various LLMs. T

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).