← all papers · overview

Conformal Tail Risk Control For Large Language Model Alignment

Abstract

Recent developments in large language models (LLMs) have led to their widespread usage for various tasks. The prevalence of LLMs in society implores the assurance on the reliability of their performance. In particular, risk-sensitive applications demand meticulous attention to unexpectedly poor outcomes, i.e., tail events, for instance, toxic answers, humiliating language, and offensive outputs. D

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).