← all papers · overview

Compromising Honesty And Harmlessness In Language Models Via Deception Attacks

Abstract

Recent research on large language models (LLMs) has demonstrated their ability to understand and employ deceptive behavior, even without explicit prompting. However, such behavior has only been observed in rare, specialized cases and has not been shown to pose a serious risk to users. Additionally, research on AI alignment has made significant advancements in training models to refuse generating m

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).