← all papers · overview

STAIR: Improving Safety Alignment With Introspective Reasoning

Abstract

Ensuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods typically suffer from safety-performance trade-offs and the susceptibility to jailbreak attacks, primarily due to their reliance on direct refusals for malicious queries. In this paper, we propose STAIR, a novel framework

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).