← all papers · overview

Following The Whispers Of Values: Unraveling Neural Mechanisms Behind Value-oriented Behaviors In Llms

Abstract

Despite the impressive performance of large language models (LLMs), they can present unintended biases and harmful behaviors driven by encoded values, emphasizing the urgent need to understand the value mechanisms behind them. However, current research primarily evaluates these values through external responses with a focus on AI safety, lacking interpretability and failing to assess social values

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).