← all papers · overview

Beyond Task Completion: Revealing Corrupt Success In LLM Agents Through Procedure-aware Evaluation

Abstract

Large Language Model (LLM)-based agents are increasingly adopted in high-stakes settings, but current benchmarks evaluate mainly whether a task was completed, not how. We introduce Procedure-Aware Evaluation (PAE), a framework that formalizes agent procedures as structured observations and exposes consistency relationships between what agents observe, communicate, and execute. PAE evaluates agents

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).