Awesome Multimodal
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Binghai Wang β most-cited papers & profile Β· Multimodal
β authors
Β·
overview
Binghai Wang
10
papers Β·
26
citations Β·
2
h-index
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
Secrets of RLHF in Large Language Models Part I: PPO
2023 Β· 19 citations
Secrets of RLHF in Large Language Models Part II: Reward Modeling
2024 Β· 7 citations
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
2026
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
2026
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
2025
WorldPM: Scaling Human Preference Modeling
2025
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
2024
Secrets of RLHF in Large Language Models Part I: PPO
2023
Secrets of RLHF in Large Language Models Part II: Reward Modeling
2024
Topics
Reinforcement Learning
Safety & Alignment
Training Techniques
Evaluation
cs.LG
cs.CL
Model-Based RL
RLHF & Alignment
cs.AI
cs.IR