← all papers · overview

When "better" Prompts Hurt: Evaluation-driven Iteration For LLM Applications

Abstract

Evaluating Large Language Model (LLM) applications differs from traditional software testing because outputs are stochastic, high-dimensional, and sensitive to prompt and model changes. We present an evaluation-driven workflow - Define, Test, Diagnose, Fix - that turns these challenges into a repeatable engineering loop. We introduce the Minimum Viable Evaluation Suite (MVES), a tiered set of re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).