← all papers · overview

MULTIVERSE: Exposing Large Language Model Alignment Problems In Diverse Worlds

Abstract

Large Language Model (LLM) alignment aims to ensure that LLM outputs match with human values. Researchers have demonstrated the severity of alignment problems with a large spectrum of jailbreak techniques that can induce LLMs to produce malicious content during conversations. Finding the corresponding jailbreaking prompts usually requires substantial human intelligence or computation resources. In

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).