← all papers · overview

WASD: Locating Critical Neurons As Sufficient Conditions For Explaining And Controlling LLM Behavior

Abstract

Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often incur high training costs, lack natural language controllability, or compromise semantic coherence. To bridge this gap, we propose WASD (unWeaving Actionable Sufficient Directives), a novel framework that explains model behavior by identifying sufficient neural condition

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).