AlpacaEval 2
Emerging12papers using it
2024first seen
'AlpacaEval 2' is a benchmark dataset used to evaluate the performance of large language models in open-domain tasks through various assessment metrics.
Papers using AlpacaEval 2 (12)
- SERL: Self-Examining Reinforcement Learning on Open-DomainReward Model Routing in AlignmentQuantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition FunctionsTGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference OptimizationDPO Meets PPO: Reinforced Token Optimization for RLHFSimPO: Simple Preference Optimization with a Reference-Free RewardRLHF Workflow: From Reward Modeling to Online RLHFThe Perfect Blend: Redefining RLHF with Mixture of JudgesBootstrapping Language Models with DPO Implicit RewardsCost-Effective Proxy Reward Model Construction with On-Policy and Active
LearningAlphaDPO: Adaptive Reward Margin for Direct Preference OptimizationT-REG: Preference Optimization with Token-Level Reward Regularization