← all papers · overview

Multiple Choice Questions: Reasoning Makes Large Language Models (llms) More Self-confident Even When They Are Wrong

Abstract

One of the most widely used methods to evaluate LLMs are Multiple Choice Question (MCQ) tests. MCQ benchmarks enable the testing of LLM knowledge on almost any topic at scale as the results can be processed automatically. To help the LLM answer, a few examples called few shots can be included in the prompt. Moreover, the LLM can be asked to answer the question directly with the selected option or

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).