← all papers · overview

Safeseek: Universal Attribution Of Safety Circuits In Language Models

Abstract

Mechanistic interpretability reveals that safety-critical behaviors (e.g., alignment, jailbreak, backdoor) in Large Language Models (LLMs) are grounded in specialized functional components. However, existing safety attribution methods struggle with generalization and reliability due to their reliance on heuristic, domain-specific metrics and search algorithms. To address this, we propose \ourmetho

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).