← all papers · overview

Tamperbench: Systematically Stress-testing LLM Safety Under Fine-tuning And Tampering

Abstract

As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks. However, there is no standard approach to evaluate tamper resistance. Varied data sets, metrics, and tampering configurations make it difficult to compare safety, utility, and robustness

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).