← all papers · overview

Evaluating Large Language Models In Ophthalmology

Abstract

Purpose: The performance of three different large language models (LLMS) (GPT-3.5, GPT-4, and PaLM2) in answering ophthalmology professional questions was evaluated and compared with that of three different professional populations (medical undergraduates, medical masters, and attending physicians). Methods: A 100-item ophthalmology single-choice test was administered to three different LLMs (GPT-

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).