Abstract
As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judicial and administrative contexts, it becomes essential to explore how they answer legal questions, and in particular the factors that lead them to decide difficult questions. A specific feature of legal decisions is the need to respond to arguments advanced by contending parties. A legal decision-maker must be able to engage with, and respond to, including through being potentially persuaded by, these arguments. Conversely, they should not be unduly persuadable, deciding cases based on the skills of the advocates rather than the merits of the case. In this paper we explore how frontier open- and closed-weights LLMs respond to legal arguments. We propose a metric to measure persuadability in the trilateral setting in which competing advocates seek to persuade a judge of opposite conclusions. We report original experimental results measuring how far the quality of the advocate making arguments affects the likelihood that a given model will agree with a particular legal point of view. We further examine how far models are capable of distinguishing between stronger and weaker arguments and how far model judgments in this domain are affected by positional bias. Through parallel bilateral trials we show how the trilateral setting changes the demands on judge models, and in turn their apparent persuadability. Finally we examine the specific features of arguments that affect persuasion, including the relative contribution of legal content and rhetorical form, the extent to which model persuasion tracks human expert judgments of argument quality, and the extent to which argument quantity, diversity and type affect persuasive outcomes. Our results have implications for the feasibility of adopting LLMs across legal and administrative settings.