← all papers · overview

Large Language Model Critics For Execution-free Evaluation Of Code Changes

Abstract

Large language models (LLMs) offer a promising way forward for automating software engineering tasks, such as bug fixes, feature additions, etc., via multi-step LLM-based agentic workflows. However, existing metrics for evaluating such workflows, mainly build status and occasionally log analysis, are too sparse and limited in providing the information needed to assess the quality of changes made.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).