← all papers · overview

Stop Listening To Me! How Multi-turn Conversations Can Degrade LLM Diagnostic Reasoning

Abstract

Patients and clinicians are increasingly using chatbots powered by large language models (LLMs) for healthcare inquiries. While state-of-the-art LLMs exhibit high performance on static diagnostic reasoning benchmarks, their efficacy across multi-turn conversations, which better reflect real-world usage, has been understudied. In this paper, we evaluate 17 LLMs across three clinical datasets to inv

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).