← all papers · overview

Towards Faithful Natural Language Explanations: A Study Using Activation Patching In Large Language Models

Abstract

Large Language Models (LLMs) are capable of generating persuasive Natural Language Explanations (NLEs) to justify their answers. However, the faithfulness of these explanations should not be readily trusted at face value. Recent studies have proposed various methods to measure the faithfulness of NLEs, typically by inserting perturbations at the explanation or feature level. We argue that these ap

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).