Awesome Reinforcement Learning
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jianfeng Gao — most-cited papers & profile · Reinforcement Learning
← authors
·
overview
Jianfeng Gao
31
papers ·
1550
citations ·
91
h-index
Intelligent Health (United Kingdom)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Deep Dyna-Q: Integrating Planning for Task-Completion Dialogue Policy Learning
2018 · 185 citations
Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
2017 · 153 citations
A User Simulator for Task-Completion Dialogues
2016 · 140 citations
BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-Oriented Dialogue Systems
2016 · 98 citations
M-Walk: Learning to Walk over Graphs using Monte Carlo Tree Search
2018 · 60 citations
Combating Reinforcement Learning's Sisyphean Curse with Intrinsic Fear
2016 · 50 citations
Subgoal Discovery for Hierarchical Dialogue Policy Learning
2018 · 39 citations
Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access
2016 · 36 citations
Guided Dialog Policy Learning without Adversarial Learning in the Loop
2020 · 13 citations
Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads
2016 · 9 citations
Switch-based Active Deep Dyna-Q: Efficient Adaptive Planning for Task-Completion Dialogue Policy Learning
2018 · 5 citations
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
2025 · 1 citations
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
2025
Top co-authors
Baolin Peng
· 5
Xiujun Li
· 5
Li Deng
· 4
Lihong Li
· 3
Zachary C. Lipton
· 3
Bhuwan Dhingra
· 2
Faisal Ahmed
· 2
Jianshu Chen
· 2
Kam‐Fai Wong
· 2
Lihong Li
· 2
Lihong Li
· 2
Yelong Shen
· 2
Topics
Model-Based RL
Value-Based
Policy Gradient
Exploration
RLHF & Alignment
Meta-RL
Multi-Agent
Safe RL
Offline RL