← all papers · overview

Trojan Activation Attack: Red-teaming Large Language Models Using Activation Steering For Safety-alignment

Abstract

To ensure AI safety, instruction-tuned Large Language Models (LLMs) are specifically trained to ensure alignment, which refers to making models behave in accordance with human intentions. While these models have demonstrated commendable results on various safety benchmarks, the vulnerability of their safety alignment has not been extensively studied. This is particularly troubling given the potent

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).