← all papers · overview

Understanding LLM Behavior When Encountering User-supplied Harmful Content In Harmless Tasks

Abstract

Large Language Models (LLMs) are increasingly trained to align with human values, primarily focusing on task level, i.e., refusing to execute directly harmful tasks. However, a subtle yet crucial content-level ethical question is often overlooked: when performing a seemingly benign task, will LLMs -- like morally conscious human beings -- refuse to proceed when encountering harmful content in user

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).