HumanEvalFix
Emerging7papers using it
2023first seen
The 'HumanEvalFix' dataset/benchmark contains a collection of coding problems designed to evaluate the ability of models to generate unit tests that effectively reveal errors in faulty code while predicting correct outputs.
Papers using HumanEvalFix (7)
- Learning to Generate Unit Tests for Automated DebuggingFrom Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical DebuggingSEAlign: Alignment Training for Software Engineering AgentSWE-agent: Agent-Computer Interfaces Enable Automated Software
EngineeringFrom Code to Correctness: Closing the Last Mile of Code Generation with
Hierarchical DebuggingCoffee: Boost Your Code LLMs by Fixing Bugs with FeedbackCode Comparison Tuning for Code Large Language Models