Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
Reward
loadingβ¦
π€
Ask AI
Awesome Reward β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
Reward
19 papers tagged Reward β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
19 papers Β· trending (default)
numbers = π₯ heat
Training Data Efficiency in Multimodal Process Reward Models
(2026)
Jinyuan Li et al.
1.94
Beyond Length Scaling: Synergizing Breadth and Depth for Generative Reward Models
(2026)
Qiyuan Zhang et al.
1.94
PRISM: Pushing the Frontier of Deep Think via Process Reward Model-Guided Inference
(2026)
Rituraj Sharma et al.
1.94
Rethinking Diverse Human Preference Learning through Principal Component Analysis
(2025)
Feng Luo et al.
1.28
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
(2025)
Yuliang Liu et al.
1.28
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
(2025)
Jingyi Zhang et al.
1.28
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
(2025)
Yibin Wang et al.
1.28
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
(2025)
Shuquan Lian et al.
1.28
Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning
(2025)
Jiuzhou Han et al.
1.28
Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
(2025)
Yong Deng et al.
1.28
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
(2025)
Chenlu Ye et al.
1.28
RewardDance: Reward Scaling in Visual Generation
(2025)
Jie Wu et al.
1.28
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
(2025)
Xin Xu et al.
1.28
Parallel Test-Time Scaling for Latent Reasoning Models
(2025)
Runyang You et al.
1.28
VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
(2025)
Qunzhong Wang et al.
1.28
Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
(2025)
Yunhong Lu et al.
1.28
Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
(2025)
Huajie Tan et al.
1.28
Efficient RLHF: Reducing the Memory Usage of PPO
(2023)
Michael Santacroce et al.
β
ICE-GRT: Instruction Context Enhancement by Generative Reinforcement based Transformers
(2024)
Chen Zheng et al.
β