← all papers · overview

An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments

Abstract

Current IR evaluation is based on relevance judgments, created either manually or automatically, with decisions outsourced to Large Language Models (LLMs). We offer an alternative paradigm, that never relies on relevance judgments in any form. Instead, a text is defined as relevant if it contains information that enables the answering of key questions. We use this idea to design the EXAM Answerabi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).