← all papers · overview

Infibench: Evaluating The Question-answering Capabilities Of Code Large Language Models

Abstract

Large Language Models for code (code LLMs) have witnessed tremendous progress in recent years. With the rapid development of code LLMs, many popular evaluation benchmarks, such as HumanEval, DS-1000, and MBPP, have emerged to measure the performance of code LLMs with a particular focus on code generation tasks. However, they are insufficient to cover the full range of expected capabilities of code

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).