← all papers · overview

Benchmarking Llama2, Mistral, Gemma And GPT For Factuality, Toxicity, Bias And Propensity For Hallucinations

Abstract

This paper introduces fourteen novel datasets for the evaluation of Large Language Models' safety in the context of enterprise tasks. A method was devised to evaluate a model's safety, as determined by its ability to follow instructions and output factual, unbiased, grounded, and appropriate content. In this research, we used OpenAI GPT as point of comparison since it excels at all levels of safet

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).