← all papers · overview

Dissociation Of Faithful And Unfaithful Reasoning In Llms

Abstract

Large language models (LLMs) often improve their performance in downstream tasks when they generate Chain of Thought reasoning text before producing an answer. We investigate how LLMs recover from errors in Chain of Thought. Through analysis of error recovery behaviors, we find evidence for unfaithfulness in Chain of Thought, which occurs when models arrive at the correct answer despite invalid re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).