← all papers · overview

Instructional Segment Embedding: Improving LLM Safety With Instruction Hierarchy

Abstract

Large Language Models (LLMs) are susceptible to security and safety threats, such as prompt injection, prompt extraction, and harmful requests. One major cause of these vulnerabilities is the lack of an instruction hierarchy. Modern LLM architectures treat all inputs equally, failing to distinguish between and prioritize various types of instructions, such as system messages, user prompts, and dat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).