← all papers · overview

Followeval: A Multi-dimensional Benchmark For Assessing The Instruction-following Capability Of Large Language Models

Abstract

The effective assessment of the instruction-following ability of large language models (LLMs) is of paramount importance. A model that cannot adhere to human instructions might be not able to provide reliable and helpful responses. In pursuit of this goal, various benchmarks have been constructed to evaluate the instruction-following capacity of these models. However, these benchmarks are limited

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).