Awesome AI Agents
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yunhao Tang — most-cited papers & profile · AI Agents
← authors
·
overview
Yunhao Tang
12
papers ·
23
citations ·
14
h-index
Beijing Institute of Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
RL-finetuning LLMs from on- and off-policy data with a single algorithm
2025 · 13 citations
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
2025 · 8 citations
Offline Regularised Reinforcement Learning for Large Language Models Alignment
2024 · 1 citations
On scalable oversight with weak LLMs judging strong LLMs
2024 · 1 citations
Ministral 3
2026
Voxtral
2025
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
2025
RL-finetuning LLMs from on- and off-policy data with a single algorithm
2025
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
2024
Understanding the performance gap between online and offline alignment algorithms
2024
Top co-authors
Aldo Pacchiano
· 1
Deepali Jain
· 1
Jack Parker-Holder
· 1
Krzysztof Choromanski
· 1
Shipra Agrawal
· 1
Tamas Sarlos
· 1
Wenbo Gao
· 1
Xingyou Song
· 1
Yuxiang Yang
· 1
Topics
Training Techniques
Safety & Alignment
Policy Gradient
Offline RL
Efficiency
Model Architecture
Vision-Language
cs.LG
Reinforcement Learning
In-Context Learning