← all papers · overview

Taxonomy-based Checklist For Large Language Model Evaluation

Abstract

As large language models (LLMs) have been used in many downstream tasks, the internal stereotypical representation may affect the fairness of the outputs. In this work, we introduce human knowledge into natural language interventions and study pre-trained language models' (LMs) behaviors within the context of gender bias. Inspired by CheckList behavioral testing, we present a checklist-style task

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).