← all papers · overview

Easy2hard-bench: Standardized Difficulty Labels For Profiling LLM Performance And Generalization

Abstract

While generalization over tasks from easy to hard is crucial to profile language models (LLMs), the datasets with fine-grained difficulty annotations for each problem across a broad range of complexity are still blank. Aiming to address this limitation, we present Easy2Hard-Bench, a consistently formatted collection of 6 benchmark datasets spanning various domains, such as mathematics and programm

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).