← all papers · overview

Detectors For Safe And Reliable Llms: Implementations, Uses, And Limitations

Abstract

Large language models (LLMs) are susceptible to a variety of risks, from non-faithful output to biased and toxic generations. Due to several limiting factors surrounding LLMs (training cost, API access, data availability, etc.), it may not always be feasible to impose direct safety constraints on a deployed model. Therefore, an efficient and reliable alternative is required. To this end, we presen

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).