← all papers · overview

Benchmarking Llm-based Relevance Judgment Methods

Abstract

Large Language Models (LLMs) are increasingly deployed in both academic and industry settings to automate the evaluation of information seeking systems, particularly by generating graded relevance judgments. Previous work on LLM-based relevance assessment has primarily focused on replicating graded human relevance judgments through various prompting strategies. However, there has been limited expl

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).