UltraFeedback

Name: UltraFeedback
License: mit

Emerging

6papers using it

5,121HF downloads

423HF likes

2025first seen

Introduction GitHub Repo UltraRM-13b UltraCM-13b UltraFeedback is a large-scale, fine-grained, diverse preference dataset, used for training powerful reward models and critic models. We collect about 64k prompts from diverse resources (including UltraChat, ShareGPT, Evol-Instruct, TruthfulQA, FalseQA, and FLAN). We the

🤗 Hugging Face⚖ mit

Papers using UltraFeedback (6)

Less is More: Improving LLM Alignment via Preference Data Selection2025 · 1 cites

MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment2026

Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs2025

When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets2025

Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information2025

CoPL: Collaborative Preference Learning for Personalizing LLMs2025