← all papers · overview

"I May Not Have Articulated Myself Clearly": Diagnosing Dynamic Instability In LLM Reasoning At Inference Time

Abstract

Reasoning failures in large language models (LLMs) are typically measured only at the end of a generation, yet many failures manifest as a process-level breakdown: the model "loses the thread" mid-reasoning. We study whether such breakdowns are detectable from inference-time observables available in standard APIs (token log probabilities), without any training or fine-tuning. We define a simple in

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).