← all papers · overview

Can Language Models Explain Their Own Classification Behavior?

Abstract

Large language models (LLMs) perform well at a myriad of tasks, but explaining the processes behind this performance is a challenge. This paper investigates whether LLMs can give faithful high-level explanations of their own internal processes. To explore this, we introduce a dataset, ArticulateRules, of few-shot text-based classification tasks generated by simple rules. Each rule is associated wi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).