Abstract
Large language models (LLMs) may generate text that lacks consistency with human knowledge, leading to factual inaccuracies or \textit\{hallucination\}. Existing research for evaluating the factuality of LLMs involves extracting fact claims using an LLM and verifying them against a predefined fact source. However, these evaluation metrics are task-specific, and not scalable, and the substitutabili