← all papers · overview

DELPHI: Data For Evaluating Llms' Performance In Handling Controversial Issues

Abstract

Controversy is a reflection of our zeitgeist, and an important aspect to any discourse. The rise of large language models (LLMs) as conversational systems has increased public reliance on these systems for answers to their various questions. Consequently, it is crucial to systematically examine how these models respond to questions that pertaining to ongoing debates. However, few such datasets exi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).