← all papers · overview

Llm-as-a-judge: Reassessing The Performance Of Llms In Extractive QA

Abstract

Extractive reading comprehension question answering (QA) datasets are typically evaluated using Exact Match (EM) and F1-score, but these metrics often fail to fully capture model performance. With the success of large language models (LLMs), they have been employed in various tasks, including serving as judges (LLM-as-a-judge). In this paper, we reassess the performance of QA models using LLM-as-a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).