← all papers · overview

Do Large Language Models Have Shared Weaknesses In Medical Question Answering?

Abstract

Large language models (LLMs) have made rapid improvement on medical benchmarks, but their unreliability remains a persistent challenge for safe real-world uses. To design for the use LLMs as a category, rather than for specific models, requires developing an understanding of shared strengths and weaknesses which appear across models. To address this challenge, we benchmark a range of top LLMs and

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).