← all papers · overview

Language Models Resist Alignment: Evidence From Data Compression

Abstract

Large language models (LLMs) may exhibit unintended or undesirable behaviors. Recent works have concentrated on aligning LLMs to mitigate harmful outputs. Despite these efforts, some anomalies indicate that even a well-conducted alignment process can be easily circumvented, whether intentionally or accidentally. Does alignment fine-tuning yield have robust effects on models, or are its impacts mer

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).