← all papers · overview

Iheval: Evaluating Language Models On Following The Instruction Hierarchy

Abstract

The instruction hierarchy, which establishes a priority order from system messages to user messages, conversation history, and tool outputs, is essential for ensuring consistent and safe behavior in language models (LMs). Despite its importance, this topic receives limited attention, and there is a lack of comprehensive benchmarks for evaluating models' ability to follow the instruction hierarchy.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).