← all papers · overview

Re-evaluating Automatic LLM System Ranking For Alignment With Human Preference

Abstract

Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consuming nature of human evaluations, an automatic LLM bencher (i.e., an automatic evaluation framework that aims to rank LLMs based on their alignment with human preferences) is indispensable. An automatic LLM bencher consist

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).