← all papers · overview

Root Defence Strategies: Ensuring Safety Of LLM At The Decoding Level

Abstract

Large language models (LLMs) have demonstrated immense utility across various industries. However, as LLMs advance, the risk of harmful outputs increases due to incorrect or malicious instruction prompts. While current methods effectively address jailbreak risks, they share common limitations: 1) Judging harmful responses from the prefill-level lacks utilization of the model's decoding outputs, le

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).