← all papers · overview

Make Them Spill The Beans! Coercive Knowledge Extraction From (production) Llms

Abstract

Large Language Models (LLMs) are now widely used in various applications, making it crucial to align their ethical standards with human values. However, recent jail-breaking methods demonstrate that this alignment can be undermined using carefully constructed prompts. In our study, we reveal a new threat to LLM alignment when a bad actor has access to the model's output logits, a common feature in

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).