← all papers · overview

LLM Reasoning Predicts When Models Are Right: Evidence From Coding Classroom Discourse

Abstract

Large Language Models (LLMs) are increasingly deployed to automatically label and analyze educational dialogue at scale, yet current pipelines lack reliable ways to detect when models are wrong. We investigate whether reasoning generated by LLMs can be used to predict the correctness of a model's own predictions. We analyze 30,300 teacher utterances from classroom dialogue, each labeled by multipl

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).