One-shot Labeling For Automatic Relevance Estimation
2023 Β· Sean MacAvaney, Luca Soldaini
Abstract
Dealing with unjudged documents ("holes") in relevance assessments is a perennial problem when evaluating search systems with offline experiments. Holes can reduce the apparent effectiveness of retrieval systems during evaluation and introduce biases in models trained with incomplete data. In this work, we explore whether large language models can help us fill such holes to improve offline evaluations. We examine an extreme, albeit common, evaluation setting wherein only a single known relevant document per query is available for evaluation. We then explore various approaches for predicting the relevance of unjudged documents with respect to a query and the known relevant document, including nearest neighbor, supervised, and prompting techniques. We find that although the predictions of these One-Shot Labelers (1SL) frequently disagree with human assessments, the labels they produce yield a far more reliable ranking of systems than the single labels do alone. Specifically, the stronges
Authors
(none)
Tags
Stats
Related papers
- Learning More From Less: Towards Strengthening Weak Supervision For Ad-hoc Retrieval (2019)5.84
- Enhancing The Ranking Context Of Dense Retrieval Methods Through Reciprocal Nearest Neighbors (2023)4.52
- Formalized Information Needs Improve Large-language-model Relevance Judgments (2026)0.00
- Precise Zero-shot Dense Retrieval Without Relevance Labels (2022)17.27
- Toward Automatic Relevance Judgment Using Vision--language Models For Image--text Retrieval Evaluation (2024)0.00
- Bixse: Improving Dense Retrieval Via Probabilistic Graded Relevance Distillation (2025)0.00
- Promptreps: Prompting Large Language Models To Generate Dense And Sparse Representations For Zero-shot Document Retrieval (2024)10.61
- The Overlooked Role Of Graded Relevance Thresholds In Multilingual Dense Retrieval (2026)0.00