← all papers · overview

View From Above: A Framework For Evaluating Distribution Shifts In Model Behavior

Abstract

When large language models (LLMs) are asked to perform certain tasks, how can we be sure that their learned representations align with reality? We propose a domain-agnostic framework for systematically evaluating distribution shifts in LLMs decision-making processes, where they are given control of mechanisms governed by pre-defined rules. While individual LLM actions may appear consistent with ex

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).