← all papers · overview

Alignment Drift In Multimodal Llms: A Two-phase, Longitudinal Evaluation Of Harm Across Eight Model Releases

Abstract

Multimodal large language models (MLLMs) are increasingly deployed in real-world systems, yet their safety under adversarial prompting remains underexplored. We present a two-phase evaluation of MLLM harmlessness using a fixed benchmark of 726 adversarial prompts authored by 26 professional red teamers. Phase 1 assessed GPT-4o, Claude Sonnet 3.5, Pixtral 12B, and Qwen VL Plus; Phase 2 evaluated th

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).