← all papers · overview

Llm-generated Black-box Explanations Can Be Adversarially Helpful

Abstract

Large Language Models (LLMs) are becoming vital tools that help us solve and understand complex problems by acting as digital assistants. LLMs can generate convincing explanations, even when only given the inputs and outputs of these problems, i.e., in a ``black-box'' approach. However, our research uncovers a hidden risk tied to this approach, which we call *adversarial helpfulness*. This happens

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).