← all papers · overview

Can We Verify Step By Step For Incorrect Answer Detection?

Abstract

Chain-of-Thought (CoT) prompting has marked a significant advancement in enhancing the reasoning capabilities of large language models (LLMs). Previous studies have developed various extensions of CoT, which focus primarily on enhancing end-task performance. In addition, there has been research on assessing the quality of reasoning chains in CoT. This raises an intriguing question: Is it possible

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).