MultiPL-E
Canonical13papers using it
2023first seen
A multi-language translation of HumanEval/MBPP for evaluating code generation across 18+ programming languages.
Papers using MultiPL-E (13)
- SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code EvaluationAgnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning EnvironmentIterative Self-Training for Code Generation via Reinforced Re-RankingReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code GenerationCode Llama: Open Foundation Models for CodeSantaCoder: don't reach for the stars!XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging
Upcycled Mixture-of-ExpertsA Preliminary Study of Multilingual Code Language Models for Code
Generation Task Using Translated BenchmarksInstruction Fusion: Advancing Prompt Evolution through HybridizationInverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct$\mathbb{USCD}$: Improving Code Generation of LLMs by Uncertainty-Aware
Selective Contrastive DecodingExecRepoBench: Multi-level Executable Code Completion EvaluationPERC: Plan-As-Query Example Retrieval for Underrepresented Code
Generation