← all papers · overview

Chain-of-thought Unfaithfulness As Disguised Accuracy

Abstract

Understanding the extent to which Chain-of-Thought (CoT) generations align with a large language model's (LLM) internal computations is critical for deciding whether to trust an LLM's output. As a proxy for CoT faithfulness, Lanham et al. (2023) propose a metric that measures a model's dependence on its CoT for producing an answer. Within a single family of proprietary models, they find that LLMs

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).