← all papers · overview

Stable Reasoning, Unstable Responses: Mitigating LLM Deception Via Stability Asymmetry

Abstract

As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical. A vital risk is intrinsic deception, wherein models strategically mislead users to achieve their own objectives. Existing alignment approaches based on chain-of-thought (CoT) monitoring supervise explicit reasoning traces. However, under optimization pressure, models are incentivized

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).