← all papers · overview

Nevermind: Instruction Override And Moderation In Large Language Models

Abstract

Given the impressive capabilities of recent Large Language Models (LLMs), we investigate and benchmark the most popular proprietary and different sized open source models on the task of explicit instruction following in conflicting situations, e.g. overrides. These include the ability of the model to override the knowledge within the weights of the model, the ability to override (or moderate) extr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).