← all papers · overview

Toolbehonest: A Multi-level Hallucination Diagnostic Benchmark For Tool-augmented Large Language Models

Abstract

Tool-augmented large language models (LLMs) are rapidly being integrated into real-world applications. Due to the lack of benchmarks, the community has yet to fully understand the hallucination issues within these models. To address this challenge, we introduce a comprehensive diagnostic benchmark, ToolBH. Specifically, we assess the LLM's hallucinations through two perspectives: depth and breadth

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).