← all papers · overview

Halluverse-m^3: A Multitask Multilingual Benchmark For Hallucination In Llms

Abstract

Hallucinations in large language models remain a persistent challenge, particularly in multilingual and generative settings where factual consistency is difficult to maintain. While recent models show strong performance on English-centric benchmarks, their behavior across languages, tasks, and hallucination types is not yet well understood. In this work, we introduce Halluverse-M^3, a dataset desi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).