Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
advantage
loadingβ¦
π€
Ask AI
Awesome advantage β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
advantage
13 papers tagged advantage β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
13 papers Β· trending (default)
numbers = π₯ heat
Reasoning with Exploration: An Entropy Perspective
(2025)
Daixuan Cheng et al.
2.87
Your Group-Relative Advantage Is Biased
(2026)
Fengkai Yang et al.
1.94
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
(2026)
Wenbo Hu et al.
1.94
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
(2026)
Shilin Yan et al.
1.94
APPO: Agentic Procedural Policy Optimization
(2026)
Xucong Wang et al.
1.94
STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
(2026)
Haipeng Luo et al.
1.94
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
(2025)
Ganqu Cui et al.
1.28
Agentic Reinforced Policy Optimization
(2025)
Guanting Dong et al.
1.28
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
(2025)
Zhipeng Chen et al.
1.28
Towards a Unified View of Large Language Model Post-Training
(2025)
Xingtai Lv et al.
1.28
GRPO-MA: Multi-Answer Generation in GRPO for Stable and Efficient Chain-of-Thought Training
(2025)
Hongcheng Wang et al.
1.28
Training-Free Group Relative Policy Optimization
(2025)
Yuzheng Cai et al.
1.28
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
(2024)
Xueliang Zhao et al.
β