← all papers · overview

The Probabilities Also Matter: A More Faithful Metric For Faithfulness Of Free-text Explanations In Large Language Models

Abstract

In order to oversee advanced AI systems, it is important to understand their underlying decision-making process. When prompted, large language models (LLMs) can provide natural language explanations or reasoning traces that sound plausible and receive high ratings from human annotators. However, it is unclear to what extent these explanations are faithful, i.e., truly capture the factors responsib

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).