UltraFeedback
Emerging6papers using it
5,121HF downloads
423HF likes
2025first seen
Introduction GitHub Repo UltraRM-13b UltraCM-13b UltraFeedback is a large-scale, fine-grained, diverse preference dataset, used for training powerful reward models and critic models. We collect about 64k prompts from diverse resources (including UltraChat, ShareGPT, Evol-Instruct, TruthfulQA, FalseQA, and FLAN). We the
π€ Hugging Faceβ mit
Papers using UltraFeedback (6)
- Less is More: Improving LLM Alignment via Preference Data SelectionMGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM AlignmentIntelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMsWhen Data is the Algorithm: A Systematic Study and Curation of Preference Optimization DatasetsBeyond Majority Voting: LLM Aggregation by Leveraging Higher-Order InformationCoPL: Collaborative Preference Learning for Personalizing LLMs