← all papers · overview

Behonest: Benchmarking Honesty In Large Language Models

Abstract

Previous works on Large Language Models (LLMs) have mainly focused on evaluating their helpfulness or harmlessness. However, honesty, another crucial alignment criterion, has received relatively less attention. Dishonest behaviors in LLMs, such as spreading misinformation and defrauding users, present severe risks that intensify as these models approach superintelligent levels. Enhancing honesty i

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).