← all papers · overview

F-eval: Assessing Fundamental Abilities With Refined Evaluation Methods

Abstract

Large language models (LLMs) garner significant attention for their unprecedented performance, leading to an increasing number of researches evaluating LLMs. However, these evaluation benchmarks are limited to assessing the instruction-following capabilities, overlooking the fundamental abilities that emerge during the pre-training stage. Previous subjective evaluation methods mainly reply on scor

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).