← all papers · overview

Building Guardrails For Large Language Models

Abstract

As Large Language Models (LLMs) become more integrated into our daily lives, it is crucial to identify and mitigate their risks, especially when the risks can have profound impacts on human users and societies. Guardrails, which filter the inputs or outputs of LLMs, have emerged as a core safeguarding technology. This position paper takes a deep look at current open-source solutions (Llama Guard,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).