← all papers · overview

MHPP: Exploring The Capabilities And Limitations Of Language Models Beyond Basic Code Generation

Abstract

Recent advancements in large language models (LLMs) have greatly improved code generation, specifically at the function level. For instance, GPT-4o has achieved a 91.0% pass rate on HumanEval. However, this draws into question the adequacy of existing benchmarks in thoroughly assessing function-level code generation capabilities. Our study analyzed two common benchmarks, HumanEval and MBPP, and fo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).