HumanEval-ET
Emerging4papers using it
2024first seen
The 'HumanEval-ET' dataset/benchmark is used to evaluate the performance of automated code generation systems, specifically focusing on their ability to generate correct Python code.
Papers using HumanEval-ET (4)
- Large Language Model Guided Self-Debugging Code GenerationModularization is Better: Effective Code Generation with Modular
PromptingCodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code GenerationSOEN-101: Code Generation by Emulating Software Process Models Using
Large Language Model Agents