← all papers · overview

Mogu: A Framework For Enhancing Safety Of Open-sourced Llms While Preserving Their Usability

Abstract

Large Language Models (LLMs) are increasingly deployed in various applications. As their usage grows, concerns regarding their safety are rising, especially in maintaining harmless responses when faced with malicious instructions. Many defense strategies have been developed to enhance the safety of LLMs. However, our research finds that existing defense strategies lead LLMs to predominantly adopt

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).