← all papers · overview

The Struggle Between Continuation And Refusal: A Mechanistic Analysis Of The Continuation-triggered Jailbreak In Llms

Abstract

With the rapid advancement of large language models (LLMs), the safety of LLMs has become a critical concern. Despite significant efforts in safety alignment, current LLMs remain vulnerable to jailbreaking attacks. However, the root causes of such vulnerabilities are still poorly understood, necessitating a rigorous investigation into jailbreak mechanisms across both academic and industrial commun

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).