← all papers · overview

When "competency" In Reasoning Opens The Door To Vulnerability: Jailbreaking Llms Via Novel Complex Ciphers

Abstract

Recent advancements in Large Language Model (LLM) safety have primarily focused on mitigating attacks crafted in natural language or common ciphers (e.g. Base64), which are likely integrated into newer models' safety training. However, we reveal a paradoxical vulnerability: as LLMs advance in reasoning, they inadvertently become more susceptible to novel jailbreaking attacks. Enhanced reasoning en

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).