Awesome AI for Code
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yanxi Chen — most-cited papers & profile · AI for Code
← authors
·
overview
Yanxi Chen
21
papers ·
137
citations ·
18
h-index
Beijing University of Chinese Medicine
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
2025 · 68 citations
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
2025 · 68 citations
Does Your ViT Still Need U-Net for Segmentation?
2026
Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
2025
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
2026
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees
2026
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees
2026
R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification
2026
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
2025
Pose State Perception of Interventional Robot for Cardio-cerebrovascular Procedures
2025
Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
2025
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
2025
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
2025
Plasma-CycleGAN: Plasma Biomarker-Guided MRI to PET Cross-modality Translation Using Conditional CycleGAN
2025
A Surface-Based Federated Chow Test Model for Integrating APOE Status, Tau Deposition Measure, and Hippocampal Surface Morphometry
2023
Topics
Reinforcement Learning
Fine-Tuning
Training Techniques
Efficiency
cs.AI
In-Context Learning
Policy Gradient
RLHF & Alignment
Offline RL
cs.LG