← all papers · overview

Foot In The Door: Understanding Large Language Model Jailbreaking Via Cognitive Psychology

Abstract

Large Language Models (LLMs) have gradually become the gateway for people to acquire new knowledge. However, attackers can break the model's security protection ("jail") to access restricted information, which is called "jailbreaking." Previous studies have shown the weakness of current LLMs when confronted with such jailbreaking attacks. Nevertheless, comprehension of the intrinsic decision-makin

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).