← all papers · overview

Evaluating LLM Safety Under Repeated Inference Via Accelerated Prompt Stress Testing

Abstract

Traditional benchmarks for large language models (LLMs) primarily assess safety risk through breadth-oriented evaluation across diverse tasks. However, real-world deployment exposes a different class of risk: operational failures arising from repeated inference on identical or near-identical prompts rather than broad task generalization. In high-stakes settings, response consistency and safety und

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).