← all papers · overview

ALLURE: Auditing And Improving Llm-based Evaluation Of Text Using Iterative In-context-learning

Abstract

From grading papers to summarizing medical documents, large language models (LLMs) are evermore used for evaluation of text generated by humans and AI alike. However, despite their extensive utility, LLMs exhibit distinct failure modes, necessitating a thorough audit and improvement of their text evaluation capabilities. Here we introduce ALLURE, a systematic approach to Auditing Large Language Mo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).