← all papers · overview

Chain-of-scrutiny: Detecting Backdoor Attacks For Large Language Models

Abstract

Large Language Models (LLMs), especially those accessed via APIs, have demonstrated impressive capabilities across various domains. However, users without technical expertise often turn to (untrustworthy) third-party services, such as prompt engineering, to enhance their LLM experience, creating vulnerabilities to adversarial threats like backdoor attacks. Backdoor-compromised LLMs generate malici

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).