How Many Parameters Does It Take To Change A Light Bulb? Evaluating Performance In Self-play Of Conversational Games As A Function Of Model Characteristics
What makes a good Large Language Model (LLM)? That it performs well on the
relevant benchmarks -- which hopefully measure, with some validity, the
presence of capabilities that are also challenged in real application. But what
makes the model perform well? What gives a model its abilities? We take a
recently introduced type of benchmark that is meant to challenge capabilities
in a goal-directed, a
Related papers
Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).