← all papers · overview

Stress-testing Capability Elicitation With Password-locked Models

Abstract

To determine the safety of large language models (LLMs), AI developers must be able to assess their dangerous capabilities. But simple prompting strategies often fail to elicit an LLM's full capabilities. One way to elicit capabilities more robustly is to fine-tune the LLM to complete the task. In this paper, we investigate the conditions under which fine-tuning-based elicitation suffices to elici

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).