← all papers · overview

Cross-modality Information Check For Detecting Jailbreaking In Multimodal Large Language Models

Abstract

Multimodal Large Language Models (MLLMs) extend the capacity of LLMs to understand multimodal information comprehensively, achieving remarkable performance in many vision-centric tasks. Despite that, recent studies have shown that these models are susceptible to jailbreak attacks, which refer to an exploitative technique where malicious users can break the safety alignment of the target model and

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).