← all papers · overview

Rlhfpoison: Reward Poisoning Attack For Reinforcement Learning With Human Feedback In Large Language Models

Abstract

Reinforcement Learning with Human Feedback (RLHF) is a methodology designed to align Large Language Models (LLMs) with human preferences, playing an important role in LLMs alignment. Despite its advantages, RLHF relies on human annotators to rank the text, which can introduce potential security vulnerabilities if any adversarial annotator (i.e., attackers) manipulates the ranking score by up-ranki

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).