← all papers · overview

On The Safety Of Open-sourced Large Language Models: Does Alignment Really Prevent Them From Being Misused?

Abstract

Large Language Models (LLMs) have achieved unprecedented performance in Natural Language Generation (NLG) tasks. However, many existing studies have shown that they could be misused to generate undesired content. In response, before releasing LLMs for public access, model developers usually align those language models through Supervised Fine-Tuning (SFT) or Reinforcement Learning with Human Feedba

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).