← all papers · overview

Online Reasoning Calibration: Test-time Training Enables Generalizable Conformal LLM Reasoning

Abstract

While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be attributed to the miscalibration of post-trained language models, and the lack of calibration in popular sampling techniques. Here, we present Online Reasoning Calibration (ORCA), a framework for calibrating the sampling p

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).