← all papers · overview

In Vino Veritas And Vulnerabilities: Examining LLM Safety Via Drunk Language Inducement

Abstract

Humans are susceptible to undesirable behaviours and privacy leaks under the influence of alcohol. This paper investigates drunk language, i.e., text written under the influence of alcohol, as a driver for safety failures in large language models (LLMs). We investigate three mechanisms for inducing drunk language in LLMs: persona-based prompting, causal fine-tuning, and reinforcement-based post-tr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).