Awesome Reinforcement Learning
Papers
Topics
Trending
Leaderboards
Authors
Datasets
Learn
Ask AI
Map
Tools
Reading Packs
News
All sections
Browse
Papers
The full index, filterable and sortable.
Topics
The same papers grouped by subject tag.
Trending
What moved this week, and by how much.
Authors
Who publishes here, and who they publish with.
Map
The collection laid out by embedding similarity.
Compare
Leaderboards
Benchmark tables, with the paper behind each number.
Datasets
The datasets these papers train and evaluate on.
Tools
Code and libraries released alongside the papers.
Read
Learn
Ordered routes from background reading to current work.
Ask AI
Ask a question and get answers cited to these papers.
Videos
The most-watched talks and lectures in this field.
Reading Packs
Short curated sets built around one question.
Follow
News
Press and coverage tied back to the papers.
Blogs
Author and lab write-ups of their own work.
Newsletter
Email digest of what changed, on a schedule.
Research Radar
Paste an abstract, get matches across every collection.
Yours
Saved
Papers you bookmarked in this browser.
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
Ask AI
Fan Yang — most-cited papers & profile · Reinforcement Learning
← authors
·
overview
Fan Yang
226
papers ·
1440
citations ·
16
h-index
Tianjin Agricultural University · Microsoft Research Asia (China)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
2025 · 206 citations
Acme: A Research Framework for Distributed Reinforcement Learning
2020 · 73 citations
DualSpec: Accelerating Deep Research Agents via Dual-Process Action Speculation
2026
ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL
2026
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning
2025
Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach
2025
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
2025
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
2025
Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models
2025
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
2025
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
2025
Representation Convergence: Mutual Distillation is Secretly a Form of Regularization
2025
Precise High-Dimensional Asymptotics for Quantifying Heterogeneous Transfers
2020
Learning Interpretable Decision Rule Sets: A Submodular Optimization Approach
2022
Improving alignment of dialogue agents via targeted human judgements
2022
Top co-authors
Bin Wen
· 3
Tingting Gao
· 3
Changyi Liu
· 2
Kaiyu Jiang
· 2
Kaiyu Tang
· 2
Mao Yang
· 2
Marco Hutter
· 2
Tianke Zhang
· 2
Wentao Zhang
· 2
Xingyu Lu
· 2
Zenan Zhou
· 2
Abbas Abdolmaleki
· 1
Topics
cs.AI
cs.LG
cs.CL
cs.CV
stat.ML
Offline RL
Multi-Agent
Value-Based
Policy Gradient