← all datasets

HumanEval

Emerging
27papers using it
2024first seen

HumanEval-X is a benchmark for the evaluation of the multilingual ability of code generative models. It consists of 820 high-quality human-crafted data samples (each with test cases) in Python, C++, Java, JavaScript, and Go, and can be used for various tasks.

Papers using HumanEval (15)

HumanEval dataset β€” papers, benchmarks & downloads Β· AI Agents