← all papers · overview

What Did I Do Wrong? Quantifying Llms' Sensitivity And Consistency To Prompt Engineering

Abstract

Large Language Models (LLMs) changed the way we design and interact with software systems. Their ability to process and extract information from text has drastically improved productivity in a number of routine tasks. Developers that want to include these models in their software stack, however, face a dreadful challenge: debugging LLMs' inconsistent behavior across minor variations of the prompt.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).