← all papers · overview

Beyond Accuracy: Evaluating Self-consistency Of Code Large Language Models With Identitychain

Abstract

Code Large Language Models (Code LLMs) are being increasingly employed in real-life applications, so evaluating them is critical. While the conventional accuracy evaluates the performance of Code LLMs on a set of individual tasks, their self-consistency across different tasks is overlooked. Intuitively, a trustworthy model should be self-consistent when generating natural language specifications f

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).