Awesome Reinforcement Learning
Papers
Topics
Trending
Leaderboards
Authors
Datasets
Learn
Ask AI
Map
Tools
Reading Packs
News
All sections
Browse
Papers
The full index, filterable and sortable.
Topics
The same papers grouped by subject tag.
Trending
What moved this week, and by how much.
Authors
Who publishes here, and who they publish with.
Map
The collection laid out by embedding similarity.
Compare
Leaderboards
Benchmark tables, with the paper behind each number.
Datasets
The datasets these papers train and evaluate on.
Tools
Code and libraries released alongside the papers.
Read
Learn
Ordered routes from background reading to current work.
Ask AI
Ask a question and get answers cited to these papers.
Videos
The most-watched talks and lectures in this field.
Reading Packs
Short curated sets built around one question.
Follow
News
Press and coverage tied back to the papers.
Blogs
Author and lab write-ups of their own work.
Newsletter
Email digest of what changed, on a schedule.
Research Radar
Paste an abstract, get matches across every collection.
Yours
Saved
Papers you bookmarked in this browser.
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
Ask AI
Pan Xu — most-cited papers & profile · Reinforcement Learning
← authors
·
overview
Pan Xu
13
papers ·
1
citations ·
43
h-index
Wuhan University · Zhongnan Hospital of Wuhan University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Policy Regularized Distributionally Robust Markov Decision Processes with Linear Function Approximation
2025 · 1 citations
Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo
2026
Sample Complexity of Distributionally Robust Off-Dynamics Reinforcement Learning with Online Interaction
2025
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
2025
Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits
2025
Linear Mixture Distributionally Robust Markov Decision Processes
2025
Finite-Time Regret of Thompson Sampling Algorithms for Exponential Family Multi-Armed Bandits
2022
Langevin Monte Carlo for Contextual Bandits
2022
Optimal Batched Best Arm Identification
2023
Finite-Time Frequentist Regret Bounds of Multi-Agent Thompson Sampling on Sparse Hypergraphs
2023
Optimal Batched Linear Bandits
2024
Top co-authors
Tianyuan Jin
· 4
Weixin Wang
· 4
Zhishuai Liu
· 3
Anima Anandkumar
· 2
Wei Deng
· 2
Xiaokui Xiao
· 2
Yiting He
· 2
Yu Yang
· 2
Eric Mazumdar
· 1
Guang Lin
· 1
Hao-Lun Hsu
· 1
Haoyang Zheng
· 1
Topics
cs.LG
stat.ML
cs.AI
cs.RO
math.ST
stat.TH
cs.CL
cs.CV
cs.MA