← all papers · overview

Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models

Abstract

Large Language Models (LLMs) often produce hallucinated or unverifiable content, undermining their reliability in factual domains. This work investigates Reinforcement Learning with Verifiable Rewards (RLVR) as a training paradigm that explicitly rewards abstention ("I don't know") alongside correctness to promote intellectual humility. We fine-tune and evaluate Granite-3.3-2B-Instruct and Qwen-3-

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).