← all papers · overview

Decoding Answers Before Chain-of-thought: Evidence From Pre-cot Probes And Activation Steering

Abstract

As chain-of-thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisions through verbalized reasoning. However, the utility of CoT toward interpretability depends upon its faithfulness -- whether the model's stated reasoning reflects the unde

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).