← all papers · overview

Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots In Ophthalmology And Llm-based Evaluation Using GPT-4

Abstract

Purpose: To assess the alignment of GPT-4-based evaluation to human clinician experts, for the evaluation of responses to ophthalmology-related patient queries generated by fine-tuned LLM chatbots. Methods: 400 ophthalmology questions and paired answers were created by ophthalmologists to represent commonly asked patient questions, divided into fine-tuning (368; 92%), and testing (40; 8%). We find

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).