← all papers · overview

From Early Encoding To Late Suppression: Interpreting Llms On Character Counting Tasks

Abstract

Large language models (LLMs) exhibit failures on elementary symbolic tasks such as character counting in a word, despite excelling on complex benchmarks. Although this limitation has been noted, the internal reasons remain unclear. We use character counting (e.g., "How many p's are in apple?") as a minimal, controlled probe that isolates token-level reasoning from higher-level confounds. Using thi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).