← all papers · overview

UFO: A Unified And Flexible Framework For Evaluating Factuality Of Large Language Models

Abstract

Large language models (LLMs) may generate text that lacks consistency with human knowledge, leading to factual inaccuracies or \textit\{hallucination\}. Existing research for evaluating the factuality of LLMs involves extracting fact claims using an LLM and verifying them against a predefined fact source. However, these evaluation metrics are task-specific, and not scalable, and the substitutabili

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).