← all papers · overview

Llms' Classification Performance Is Overclaimed

Abstract

In many classification tasks designed for AI or human to solve, gold labels are typically included within the label space by default, often posed as "which of the following is correct?" This standard setup has traditionally highlighted the strong performance of advanced AI, particularly top-performing Large Language Models (LLMs), in routine classification tasks. However, when the gold label is in

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).