← all papers · overview

MACPO: Weak-to-strong Alignment Via Multi-agent Contrastive Preference Optimization

Abstract

As large language models (LLMs) are rapidly advancing and achieving near-human capabilities on specific tasks, aligning them with human values is becoming more urgent. In scenarios where LLMs outperform humans, we face a weak-to-strong alignment problem where we need to effectively align strong student LLMs through weak supervision generated by weak teachers. Existing alignment methods mainly focu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).