← all papers · overview

Fragile Reasoning: A Mechanistic Analysis Of LLM Sensitivity To Meaning-preserving Perturbations

Abstract

Large language models demonstrate strong performance on mathematical reasoning benchmarks, yet remain surprisingly fragile to meaning-preserving surface perturbations. We systematically evaluate three open-weight LLMs, Mistral-7B, Llama-3-8B, and Qwen2.5-7B, on 677 GSM8K problems paired with semantically equivalent variants generated through name substitution and number format paraphrasing. All th

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).