← all papers · overview

Codeattack: Revealing Safety Generalization Challenges Of Large Language Models Via Code Completion

Abstract

The rapid advancement of Large Language Models (LLMs) has brought about remarkable generative capabilities but also raised concerns about their potential misuse. While strategies like supervised fine-tuning and reinforcement learning from human feedback have enhanced their safety, these methods primarily focus on natural languages, which may not generalize to other domains. This paper introduces C

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).