← all papers · overview

Pruning For Protection: Increasing Jailbreak Resistance In Aligned Llms Without Fine-tuning

Abstract

This paper investigates the impact of model compression on the way Large Language Models (LLMs) process prompts, particularly concerning jailbreak resistance. We show that moderate WANDA pruning can enhance resistance to jailbreaking attacks without fine-tuning, while maintaining performance on standard benchmarks. To systematically evaluate this safety enhancement, we introduce a dataset of 225 h

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).