← all papers · overview

Mind The Inconspicuous: Revealing The Hidden Weakness In Aligned Llms' Refusal Boundaries

Abstract

Recent advances in Large Language Models (LLMs) have led to impressive alignment where models learn to distinguish harmful from harmless queries through supervised finetuning (SFT) and reinforcement learning from human feedback (RLHF). In this paper, we reveal a subtle yet impactful weakness in these aligned models. We find that simply appending multiple end of sequence (eos) tokens can cause a ph

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).