← all papers · overview

Adversarial Representation Engineering: A General Model Editing Framework For Large Language Models

Abstract

Since the rapid development of Large Language Models (LLMs) has achieved remarkable success, understanding and rectifying their internal complex mechanisms has become an urgent issue. Recent research has attempted to interpret their behaviors through the lens of inner representation. However, developing practical and efficient methods for applying these representations for general and flexible mod

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).