← all papers · overview

Do Large Language Models Perform The Way People Expect? Measuring The Human Generalization Function

Abstract

What makes large language models (LLMs) impressive is also what makes them hard to evaluate: their diversity of uses. To evaluate these models, we must understand the purposes they will be used for. We consider a setting where these deployment decisions are made by people, and in particular, people's beliefs about where an LLM will perform well. We model such beliefs as the consequence of a human

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).