← all papers · overview

Evaluating LLM Alignment With Human Trust Models

Abstract

Trust plays a pivotal role in enabling effective cooperation, reducing uncertainty, and guiding decision-making in both human interactions and multi-agent systems. Although it is significant, there is limited understanding of how large language models (LLMs) internally conceptualize and reason about trust. This work presents a white-box analysis of trust representation in EleutherAI/gpt-j-6B, usin

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).