Awesome Reinforcement Learning
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jingren Zhou — most-cited papers & profile · Reinforcement Learning
← authors
·
overview
Jingren Zhou
106
papers ·
1122
citations ·
0
h-index
Alibaba Group (China)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Group Sequence Policy Optimization
2025 · 453 citations
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
2025 · 68 citations
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
2024 · 13 citations
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
2026
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
2026
Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
2025
Experience Augmented Policy Optimization for LLM Reasoning
2026
One-Way Policy Optimization for Self-Evolving LLMs
2026
FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
2026
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
2026
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
2026
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
2026
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
2025
AgentEvolver: Towards Efficient Self-Evolving Agent System
2025
Soft Adaptive Policy Optimization
2025
Top co-authors
Bolin Ding
· 6
Guoyin Wang
· 6
An Yang
· 5
Bowen Yu
· 5
Chiyu Ma
· 5
Junyang Lin
· 5
Kexin Huang
· 5
Jinda Lu
· 4
Shuo Yang
· 4
Chujie Zheng
· 3
Fei Huang
· 3
Haoming Meng
· 3
Topics
cs.LG
cs.AI
cs.CL
Policy Gradient
Model-Based RL
RLHF & Alignment
Offline RL
Meta-RL
Multi-Agent
Value-Based