← all papers · overview

Bias And Volatility: A Statistical Framework For Evaluating Large Language Model's Stereotypes And The Associated Generation Inconsistency

Abstract

We present a novel statistical framework for analyzing stereotypes in large language models (LLMs) by systematically estimating the bias and variation in their generation. Current alignment evaluation metrics often overlook stereotypes' randomness caused by LLMs' inconsistent generative behavior. For instance, LLMs may display contradictory stereotypes, such as those related to gender or race, for

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).