← all papers · overview

Understanding And Mitigating Gender Bias In Llms Via Interpretable Neuron Editing

Abstract

Large language models (LLMs) often exhibit gender bias, posing challenges for their safe deployment. Existing methods to mitigate bias lack a comprehensive understanding of its mechanisms or compromise the model's core capabilities. To address these issues, we propose the CommonWords dataset, to systematically evaluate gender bias in LLMs. Our analysis reveals pervasive bias across models and iden

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).