We present a method for systematically evaluating the correctness and
robustness of instruction-tuned large language models (LLMs) for code
generation via a new benchmark, Turbulence. Turbulence consists of a large set
of natural language {questiontemplates}, each of which is a
programming problem, parameterised so that it can be asked in many different
forms. Each question template
Related papers
Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).