← all papers · overview

Rethinking The Value Of Agent-generated Tests For Llm-based Software Engineering Agents

Abstract

Large Language Model (LLM) code agents increasingly resolve repository-level issues by iteratively editing code, invoking tools, and validating candidate patches. In these workflows, agents often write tests on the fly, but the value of this behavior remains unclear. For example, GPT-5.2 writes almost no new tests yet achieves performance comparable to top-ranking agents.This raises a central ques

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).