← all papers · overview

Expguard: LLM Content Moderation In Specialized Domains

Abstract

With the growing deployment of large language models (LLMs) in real-world applications, establishing robust safety guardrails to moderate their inputs and outputs has become essential to ensure adherence to safety policies. Current guardrail models predominantly address general human-LLM interactions, rendering LLMs vulnerable to harmful and adversarial content within domain-specific contexts, par

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).