DevEval
Emerging7papers using it
2024first seen
The 'DevEval' dataset/benchmark is used to evaluate the performance of code generation models by assessing their ability to handle repository-level code with cross-file dependencies and structural context.
Papers using DevEval (7)
- Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and BeyondFrom Context to Intent: Reasoning-Guided Function-Level Code CompletionAdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code GenerationCan Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'DevEval: A Manually-Annotated Code Generation Benchmark Aligned with
Real-World Code RepositoriesDevEval: Evaluating Code Generation in Practical Software ProjectsPrompting Large Language Models to Tackle the Full Software Development
Lifecycle: A Case Study