MBPP
Canonical140papers using it
168,974HF downloads
233HF likes
2021first seen
Mostly Basic Python Problems โ ~1,000 entry-level programming tasks with tests, for evaluating code generation.
๐ค Hugging Faceโ cc-by-4.0
Papers using MBPP (140)
- A Survey on Large Language Models for Code GenerationCODESIM: Multi-Agent Code Generation and Problem Solving through
Simulation-Driven Planning and DebuggingKodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for
CodingKodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for CodingRegression Accumulation in Multi-Turn LLM Programming ConversationsType-Constrained Code Generation with Language ModelsGrammar-Based Code Representation: Is It a Worthy Pursuit for LLMs?Poison with Style: A Practical Poisoning Attack on Code Large Language ModelsFocused-DPO: Enhancing Code Generation Through Focused Preference Optimization on Error-Prone PointsEnhancing LLM-Based Code Generation with Complexity Metrics: A Feedback-Driven ApproachDynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code GenerationLearning to Generate Unit Tests for Automated DebuggingQualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM
Quality ChecksAutomated Repair of Ambiguous Problem Descriptions for LLM-Based Code GenerationOpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMsCODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and DebuggingCodeMirage: Hallucinations in Code Generated by Large Language ModelsSePO: Self-Evolving Prompt Agent for System Prompt OptimizationUsing Semantic Distance to Estimate Uncertainty in LLM-Based Code GenerationAn Execution-Verified Multi-Language Benchmark for Code Semantic ReasoningACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference OptimizationPrompt Optimization for LLM Code Generation via Reinforcement LearningDomain-Adaptable Reinforcement Learning for Code Generation with Dense RewardsImproving Small Language Models for Code Generation with Reinforcement Learning from Verification FeedbackEvaluating the Environmental Impact of using SLMs and Prompt Engineering for Code GenerationFLeX: Fourier-based Low-rank EXpansion for multilingual transferMIST-RL: Mutation-based Incremental Suite Testing via Reinforcement LearningReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement LearningThink Anywhere in Code GenerationTextBFGS: A Case-Based Reasoning Approach to Code Optimization via Error-Operator RetrievalBatCoder: Self-Supervised Bidirectional Code-Documentation Learning via Back-TranslationAssessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical StudyNOIR: Privacy-Preserving Generation of Code with Open-Source LLMsMulti-task Code LLMs: Data Mix or Model Merge?Adaptive Confidence Gating in Multi-Agent Collaboration for Efficient and Optimized Code GenerationPay for Hints, Not Answers: LLM Shepherding for Cost-Efficient InferenceShape of Thought: When Distribution Matters More than Correctness in Reasoning TasksSummary-Mediated Repair: Can LLMs use code summarisation as a tool for program repair?ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code GenerationAdapTrack: Constrained Decoding without Distorting LLM's Output IntentSALT4Decompile: Inferring Source-level Abstract Logic Tree for LLM-Based Binary DecompilationProgram Synthesis via Test-Time TransductionReinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code GenerationProtocode: Prototype-Driven Interpretability for Code Generation in LLMsStatic Analysis as a Feedback Loop: Enhancing LLM-Generated Code Beyond CorrectnessAlignment with Fill-In-the-Middle for Enhancing Code GenerationSTEPWISE-CODEX-Bench: Evaluating Complex Multi-Function Comprehension and Fine-Grained Execution ReasoningAssessing Small Language Models for Code Generation: An Empirical Study with BenchmarksDr. Boot: Bootstrapping Program Synthesis Language Models to Perform RepairingCREME: Robustness Enhancement of Code LLMs via Layer-Aware Model EditingMemoCoder: Automated Function Synthesis using LLM-Supported AgentsEfficient Code LLM Training via Distribution-Consistent and Diversity-Aware Data SelectionWhen Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task DescriptionsAdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code GenerationCodeMixBench: Evaluating Large Language Models on Code Generation with
Code-Mixed PromptsWeb-Bench: A LLM Code Benchmark Based on Web Standards and FrameworksRethinking Repetition Problems of LLMs in Code GenerationSelf-Correcting Code Generation Using Small Language ModelsWhat I cannot execute, I do not understand: Training and Evaluating LLMs
on Program Execution TracesModularization is Better: Effective Code Generation with Modular
PromptingMemorize or Generalize? Evaluating LLM Code Generation with Code RewritingACECODER: Acing Coder RL via Automated Test-Case SynthesisReasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language
Models Through Logic Unit AlignmentUnitCoder: Scalable Iterative Code Synthesis with Unit Test GuidanceLearning to Solve and Verify: A Self-Play Framework for Code and Test GenerationCodeCriticBench: A Holistic Code Critique Benchmark for Large Language
ModelsThinking Before Running! Efficient Code Generation with Thorough Exploration and Optimal RefinementIsolating Language-Coding from Problem-Solving: Benchmarking LLMs with
PseudoEvalPragmatic Reasoning improves LLM Code GenerationCodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code
GenerationPlanning-Driven Programming: A Large Language Model Programming WorkflowContext-Augmented Code Generation Using Programming Knowledge GraphsCan Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'EffiLearner: Enhancing Efficiency of Generated Code via Self-OptimizationReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code GenerationCode Llama: Open Foundation Models for CodeWizardCoder: Empowering Code Large Language Models with Evol-InstructTeaching Large Language Models to Self-DebugCodeT: Code Generation with Generated TestsThe Stack: 3 TB of permissively licensed source codeDebug like a Human: A Large Language Model Debugger via Verifying
Runtime Execution Step-by-stepCYCLE: Learning to Self-Refine the Code GenerationProgram Synthesis with Large Language ModelsLiveCodeBench: Holistic and Contamination Free Evaluation of Large
Language Models for CodeTowards AI-Assisted Synthesis of Verified Dafny MethodsInteractive Code Generation via Test-Driven User-Intent FormalizationPrompt Engineering or Fine-Tuning: An Empirical Assessment of LLMs for
CodeMapCoder: Multi-Agent Code Generation for Competitive Problem SolvingOn Evaluating the Efficiency of Source Code Generated by LLMsAgentCoder: Multi-Agent-based Code Generation with Iterative Testing and
OptimisationStructured Chain-of-Thought Prompting for Code GenerationAceCoder: Utilizing Existing Code to Enhance Code GenerationFault-Aware Neural Code RankersMultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural
Code GenerationCrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code
CompletionSOEN-101: Code Generation by Emulating Software Process Models Using
Large Language Model AgentsPythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMsInterCode: Standardizing and Benchmarking Interactive Coding with
Execution FeedbackLarge Language Model-Aware In-Context Learning for Code GenerationOOP: Object-Oriented Programming Evaluation Benchmark for Large Language
ModelsReCode: Robustness Evaluation of Code Generation ModelsNExT: Teaching Large Language Models to Reason about Code ExecutionOpenCodeInterpreter: Integrating Code Generation with Execution and
RefinementSoftware Vulnerability and Functionality Assessment using LLMsLeTI: Learning to Generate from Textual InteractionsEnhancing Large Language Models in Coding Through Multi-Perspective
Self-ConsistencyCodeChain: Towards Modular Code Generation Through Chain of
Self-revisions with Representative Sub-modulesXFT: Unlocking the Power of Code Instruction Tuning by Simply Merging
Upcycled Mixture-of-ExpertsNaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and
Natural User PromptsMHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code GenerationRLTF: Reinforcement Learning from Unit Test FeedbackThe Program Testing Ability of Large Language Models for CodeDivide-and-Conquer Meets Consensus: Unleashing the Power of Functions in
Code GenerationPlanning In Natural Language Improves LLM Search For Code GenerationSelection of Prompt Engineering Techniques for Code Generation through
Predicting Code ComplexityTraining Language Models on Synthetic Edit Sequences Improves Code
SynthesisCodeTree: Agent-guided Tree Search for Code Generation with Large
Language ModelsPerfCodeGen: Improving Performance of LLM Generated Code with Execution
FeedbackAlphaVerus: Bootstrapping Formally Verified Code Generation through
Self-Improving Translation and TreefinementDecoding Data Quality via Synthetic Corruptions: Embedding-guided
Pruning of Code DataInstruction Fusion: Advancing Prompt Evolution through HybridizationUnsupervised Evaluation of Code LLMs with Round-Trip CorrectnessDolphCoder: Echo-Locating Code Large Language Models with Diverse and
Multi-Objective Instruction TuningTest-Driven Development for Code GenerationComments as Natural Logic Pivots: Improve Code Generation via Comment
PerspectiveInfiBench: Evaluating the Question-Answering Capabilities of Code Large
Language ModelsUncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via
Code Rewriting$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy
Preference Learning Driven By Synthetic Test CasesCode-Optimise: Self-Generated Preference Data for Correctness and
EfficiencyBrevity is the soul of wit: Pruning long files for code generationInverseCoder: Self-improving Instruction-Tuned Code LLMs with
Inverse-InstructBridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer$\mathbb{USCD}$: Improving Code Generation of LLMs by Uncertainty-Aware
Selective Contrastive DecodingSelf-Explained Keywords Empower Large Language Models for Code
GenerationFALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization systemDemo-Craft: Using In-Context Learning to Improve Code Generation in
Large Language ModelsDSTC: Direct Preference Learning with Only Self-Generated Tests and Code
to Improve Code LMsHumanEval Pro and MBPP Pro: Evaluating Large Language Models on
Self-invoking Code GenerationPolicy Filtration in RLHF to Fine-Tune LLM for Code GenerationCan Language Models Replace Programmers? REPOCOD Says 'Not Yet'