← all papers · overview

Efficient Llm-jailbreaking Via Multimodal-llm Jailbreak

Abstract

This paper focuses on jailbreaking attacks against large language models (LLMs), eliciting them to generate objectionable content in response to harmful user queries. Unlike previous LLM-jailbreak methods that directly orient to LLMs, our approach begins by constructing a multimodal large language model (MLLM) built upon the target LLM. Subsequently, we perform an efficient MLLM jailbreak and obta

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).