CRUXEval
Canonical10papers using it
2024first seen
CRUXEval: Code Reasoning, Understanding, and Execution Evaluation π Home Page β’ π» GitHub Repository β’ π Leaderboard β’ π Sample Explorer CRUXEval (Code Reasoning, Understanding, and eXecution Evaluation) is a benchmark of 800 Python functions and input-output pairs. The benchmark consists of two tasks, CRUXEval-I (i
Papers using CRUXEval (10)
- A Tool for In-depth Analysis of Code Execution Reasoning of Large
Language ModelsSpecEval: Evaluating Code Comprehension in Large Language Models via Program SpecificationsStepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement LearningHow Robustly do LLMs Understand Execution Semantics?Towards a Neural Debugger for PythonGenerating Verifiable Chain of Thoughts from Exection-TracesSTEPWISE-CODEX-Bench: Evaluating Complex Multi-Function Comprehension and Fine-Grained Execution ReasoningAre Large Language Models Robust in Understanding Code Against Semantics-Preserving Mutations?What I cannot execute, I do not understand: Training and Evaluating LLMs on Program Execution TracesCRUXEval: A Benchmark for Code Reasoning, Understanding and Execution