← all papers · overview

Can Large Language Models Self-correct In Medical Question Answering? An Exploratory Study

Abstract

Large language models (LLMs) have achieved strong performance on medical question answering (medical QA), and chain-of-thought (CoT) prompting has further improved results by eliciting explicit intermediate reasoning; meanwhile, self-reflective (self-corrective) prompting has been widely claimed to enhance model reliability by prompting LLMs to critique and revise their own reasoning, yet its effe

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).