← all papers · overview

Efficient Model-agnostic Alignment Via Bayesian Persuasion

Abstract

With recent advancements in large language models (LLMs), alignment has emerged as an effective technique for keeping LLMs consensus with human intent. Current methods primarily involve direct training through Supervised Fine-tuning (SFT) or Reinforcement Learning from Human Feedback (RLHF), both of which require substantial computational resources and extensive ground truth data. This paper explo

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).