← all datasets

UltraFeedback

Emerging
5papers using it
7,089HF downloads
430HF likes
2024first seen

Introduction GitHub Repo UltraRM-13b UltraCM-13b UltraFeedback is a large-scale, fine-grained, diverse preference dataset, used for training powerful reward models and critic models. We collect about 64k prompts from diverse resources (including UltraChat, ShareGPT, Evol-Instruct, TruthfulQA, FalseQA, and FLAN). We the

Papers using UltraFeedback (5)

UltraFeedback dataset β€” papers, benchmarks & downloads Β· Reinforcement Learning