← all papers · overview

Protecting Your Llms With Information Bottleneck

Abstract

The advent of large language models (LLMs) has revolutionized the field of natural language processing, yet they might be attacked to produce harmful content. Despite efforts to ethically align LLMs, these are often fragile and can be circumvented by jailbreaking attacks through optimized or manual adversarial prompts. To address this, we introduce the Information Bottleneck Protector (IBProtector

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).