Awesome Reinforcement Learning
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Baolin Peng — most-cited papers & profile · Reinforcement Learning
← authors
·
overview
Baolin Peng
26
papers ·
810
citations ·
32
h-index
Zhejiang University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Deep Dyna-Q: Integrating Planning for Task-Completion Dialogue Policy Learning
2018 · 185 citations
Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
2017 · 153 citations
Guided Dialog Policy Learning without Adversarial Learning in the Loop
2020 · 13 citations
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
2025 · 1 citations
Hierarchical Reinforcement Learning for Automatic Disease Diagnosis
2020 · 1 citations
The Trickle-down Impact of Reward (In-)consistency on RLHF
2023 · 1 citations
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
2025
ThetaEvolve: Test-time Learning on Open Problems
2025
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
2025
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
2025
Top co-authors
Jianfeng Gao
· 5
Hao Cheng
· 2
Janardhan Kulkarni
· 2
Kam‐Fai Wong
· 2
Liliang Ren
· 2
Qianhui Wu
· 2
Simon Shaolei Du
· 2
Weizhu Chen
· 2
Xiujun Li
· 2
Yelong Shen
· 2
Zhiyuan Zeng
· 2
Andrea Tupini
· 1
Topics
Model-Based RL
Policy Gradient
Exploration
RLHF & Alignment
Value-Based
Meta-RL
Offline RL
Multi-Agent
Safe RL