EvalPlus
Emerging8papers using it
2024first seen
Evalplus is a benchmark dataset used to evaluate the performance of code generation models, specifically focusing on their ability to generate code while preserving privacy and security.
Papers using EvalPlus (8)
- LLM-Powered Test Case Generation for Detecting Bugs in Plausible ProgramsUnify and Triumph: Polyglot, Diverse, and Self-Consistent Generation of
Unit Tests with LLMsBeyond Translation Accuracy: Addressing False Failures in LLM-Based Code TranslationNOIR: Privacy-Preserving Generation of Code with Open-Source LLMsOpenCodeInterpreter: Integrating Code Generation with Execution and RefinementLow-Cost Language Models: Survey and Performance Evaluation on Python
Code GenerationTowards Large Language Model Aided Program RefinementBeyond Code Generation: Assessing Code LLM Maturity with Postconditions