← all papers · overview

Answering The Wrong Question: Reasoning Trace Inversion For Abstention In Llms

Abstract

For Large Language Models (LLMs) to be reliably deployed, models must effectively know when not to answer: abstain. Reasoning models, in particular, have gained attention for impressive performance on complex tasks. However, reasoning models have been shown to have worse abstention abilities. Taking the vulnerabilities of reasoning models into account, we propose our Query Misalignment Framework.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).