← all papers · overview

Codeif-bench: Evaluating Instruction-following Capabilities Of Large Language Models In Interactive Code Generation

Abstract

Large Language Models (LLMs) have demonstrated exceptional performance in code generation tasks and have become indispensable programming assistants for developers. However, existing code generation benchmarks primarily assess the functional correctness of code generated by LLMs in single-turn interactions. They offer limited insight into LLMs' abilities to generate code that strictly follows user

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).