← all papers · overview

Turbulence: Systematically And Automatically Testing Instruction-tuned Large Language Models For Code

Abstract

We present a method for systematically evaluating the correctness and robustness of instruction-tuned large language models (LLMs) for code generation via a new benchmark, Turbulence. Turbulence consists of a large set of natural language , each of which is a programming problem, parameterised so that it can be asked in many different forms. Each question template

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).