← all papers · overview

Pretrainrl: Alleviating Factuality Hallucination Of Large Language Models At The Beginning

Abstract

Large language models (LLMs), despite their powerful capabilities, suffer from factual hallucinations where they generate verifiable falsehoods. We identify a root of this issue: the imbalanced data distribution in the pretraining corpus, which leads to a state of "low-probability truth" and "high-probability falsehood". Recent approaches, such as teaching models to say "I don't know" or post-hoc

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).