← all papers · overview

Summon A Demon And Bind It: A Grounded Theory Of LLM Red Teaming

Abstract

Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition of how and why people perform such attacks, defining LLM red-teaming based on extensive and diverse evidence. Using a formal qualitative methodology, we interviewed dozens of practitioners from a broad range of background

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).