← all papers · overview

Large Language Models As Students Who Think Aloud: Overly Coherent, Verbose, And Confident

Abstract

Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive judgments? Existing evaluations emphasize problem-solving accuracy, overlooking the fragmented and imperfect reasoning that characterizes human learning. We evaluate LLMs as novices using 630 think-aloud utterances from multi-step chemistry tutoring prob

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).