← all papers · overview

Prompt Packer: Deceiving Llms Through Compositional Instruction With Hidden Attacks

Abstract

Recently, Large language models (LLMs) with powerful general capabilities have been increasingly integrated into various Web applications, while undergoing alignment training to ensure that the generated content aligns with user intent and ethics. Unfortunately, they remain the risk of generating harmful content like hate speech and criminal activities in practical applications. Current approaches

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).