← all papers · overview

Detecting And Correcting Reference Hallucinations In Commercial Llms And Deep Research Agents

Abstract

Large language models and deep research agents supply citation URLs to support their claims, yet the reliability of these citations has not been systematically measured. We address six research questions about citation URL validity using 10 models and agents on DRBench (53,090 URLs) and 3 models on ExpertQA (168,021 URLs across 32 academic fields). We find that 3--13% of citation URLs are hallucin

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).