← all papers · overview

An Automatic Evaluation Framework For Multi-turn Medical Consultations Capabilities Of Large Language Models

Abstract

Large language models (LLMs) have achieved significant success in interacting with human. However, recent studies have revealed that these models often suffer from hallucinations, leading to overly confident but incorrect judgments. This limits their application in the medical domain, where tasks require the utmost accuracy. This paper introduces an automated evaluation framework that assesses the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).