← all papers · overview

How Many Parameters Does It Take To Change A Light Bulb? Evaluating Performance In Self-play Of Conversational Games As A Function Of Model Characteristics

Abstract

What makes a good Large Language Model (LLM)? That it performs well on the relevant benchmarks -- which hopefully measure, with some validity, the presence of capabilities that are also challenged in real application. But what makes the model perform well? What gives a model its abilities? We take a recently introduced type of benchmark that is meant to challenge capabilities in a goal-directed, a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).