Abstract
Value alignment is central to the development of safe and socially compatible artificial intelligence. However, how Large Language Models (LLMs) represent and enact human values in real-world decision contexts remains under-explored. We present ValAct-15k, a dataset of 3,000 advice-seeking scenarios derived from Reddit, designed to elicit ten values defined by Schwartz Theory of Basic Human Values. Using both the scenario-based questions and the traditional value questionnaire, we evaluate ten frontier LLMs (five from U.S. companies, five from Chinese ones) and human participants (). We find near-perfect cross-model consistency in scenario-based decisions (Pearson ), contrasting sharply with the broad variability observed among humans (). Yet, both humans and LLMs show weak correspondence between self-reported and enacted values (), revealing a systematic knowledge-action gap. When instructed to "hold" a specific value, LLMs' performance declines up to compared to merely selecting the value, indicating a role-play aversion. These findings suggest that while alignment training yields normative value convergence, it does not eliminate the human-like incoherence between knowing and acting upon values.