← all papers · overview

RES-Q: Evaluating Code-editing Large Language Model Systems At The Repository Scale

Abstract

The instruction-following ability of Large Language Models (LLMs) has cultivated a class of LLM-based systems capable of approaching complex tasks such as making edits to large code repositories. Due to the high sensitivity and unpredictability of LLM behavior in response to changes in prompting, robust evaluation tools are needed to drive future iteration of these systems. We propose RES-Q, a nat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).