← all papers · overview

Fragile Thoughts: How Large Language Models Handle Chain-of-thought Perturbations

Abstract

Chain-of-Thought (CoT) prompting has emerged as a foundational technique for eliciting reasoning from Large Language Models (LLMs), yet the robustness of this approach to corruptions in intermediate reasoning steps remains poorly understood. This paper presents a comprehensive empirical evaluation of LLM robustness to a structured taxonomy of 5 CoT perturbation types: \textit\{MathError, UnitConve

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).