← all papers · overview

A Causal Lens For Evaluating Faithfulness Metrics

Abstract

Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability. However, despite their plausibility, they may not reflect the model's true reasoning faithfully. While several faithfulness metrics have been proposed, they are often evaluated in isolation, making principled comparisons between them difficult. We present Ca

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).