← all papers · overview

A Proposed S.C.O.R.E. Evaluation Framework For Large Language Models : Safety, Consensus, Objectivity, Reproducibility And Explainability

Abstract

A comprehensive qualitative evaluation framework for large language models (LLM) in healthcare that expands beyond traditional accuracy and quantitative metrics needed. We propose 5 key aspects for evaluation of LLMs: Safety, Consensus, Objectivity, Reproducibility and Explainability (S.C.O.R.E.). We suggest that S.C.O.R.E. may form the basis for an evaluation framework for future LLM-based models

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).