← all papers · overview

Improving LLM Reliability Through Hybrid Abstention And Adaptive Detection

Abstract

Large Language Models (LLMs) deployed in production environments face a fundamental safety-utility trade-off either a strict filtering mechanisms prevent harmful outputs but often block benign queries or a relaxed controls risk unsafe content generation. Conventional guardrails based on static rules or fixed confidence thresholds are typically context-insensitive and computationally expensive, res

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).