HumanEval
Emerging3papers using it
2025first seen
The 'HumanEval' dataset is a benchmark that contains programming problems used to evaluate the performance of large reasoning models in generating code solutions.
The 'HumanEval' dataset is a benchmark that contains programming problems used to evaluate the performance of large reasoning models in generating code solutions.