← all papers · overview

Liecraft: A Multi-agent Framework For Evaluating Deceptive Capabilities In Language Models

Abstract

Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work, we present LieCraft: a novel evaluation framework and sandbox for measuring LLM deception that addresses key limitations of prior game-based evaluations. At its

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).