← all papers · overview

Evaluating Gender Bias In Large Language Models Via Chain-of-thought Prompting

Abstract

There exist both scalable tasks, like reading comprehension and fact-checking, where model performance improves with model size, and unscalable tasks, like arithmetic reasoning and symbolic reasoning, where model performance does not necessarily improve with model size. Large language models (LLMs) equipped with Chain-of-Thought (CoT) prompting are able to make accurate incremental predictions eve

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).