← all papers · overview

Evaluating Llm-based Test Generation Under Software Evolution

Abstract

Large Language Models (LLMs) are increasingly used for automated unit test generation. However, it remains unclear whether these tests reflect genuine reasoning about program behavior or simply reproduce superficial patterns learned during training. If the latter dominates, LLM-generated tests may exhibit weaknesses such as reduced coverage, missed regressions, and undetected faults. Understanding

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).