← all papers · overview

Multi-turn Evaluation Of Anthropomorphic Behaviours In Large Language Models

Abstract

The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for empirically evaluating anthropomorphic LLM behaviours in realistic and varied settings. Going beyond single-turn static benchmarks, we contribute three methodological advances in state-of-the-art (SOTA) LLM evaluation. F

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).