← all papers · overview

Sage: Evaluating Moral Consistency In Large Language Models

Abstract

Despite recent advancements showcasing the impressive capabilities of Large Language Models (LLMs) in conversational systems, we show that even state-of-the-art LLMs are morally inconsistent in their generations, questioning their reliability (and trustworthiness in general). Prior works in LLM evaluation focus on developing ground-truth data to measure accuracy on specific tasks. However, for mor

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).