Abstract
We consider the inference for the ranking of large language models (LLMs). Alignment arises as a significant challenge to mitigate hallucinations in the use of LLMs. Ranking LLMs has proven to be an effective tool to improve alignment based on the best-of- policy. In this paper, we propose a new inferential framework for hypothesis testing among the ranking for language models. Our framework