← all papers · overview

Erbench: An Entity-relationship Based Automatically Verifiable Hallucination Benchmark For Large Language Models

Abstract

Large language models (LLMs) have achieved unprecedented performances in various applications, yet evaluating them is still challenging. Existing benchmarks are either manually constructed or are automatic, but lack the ability to evaluate the thought process of LLMs with arbitrary complexity. We contend that utilizing existing relational databases based on the entity-relationship (ER) model is a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).