← all papers · overview

Imposter.ai: Adversarial Attacks With Hidden Intentions Towards Aligned Large Language Models

Abstract

With the development of large language models (LLMs) like ChatGPT, both their vast applications and potential vulnerabilities have come to the forefront. While developers have integrated multiple safety mechanisms to mitigate their misuse, a risk remains, particularly when models encounter adversarial inputs. This study unveils an attack mechanism that capitalizes on human conversation strategies

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).