Awesome AI for Code
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Ximing Lu — most-cited papers & profile · AI for Code
← authors
·
overview
Ximing Lu
30
papers ·
89
citations ·
17
h-index
University of Washington
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Quark: Controllable Text Generation with Reinforced Unlearning
2022 · 45 citations
The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
2023 · 10 citations
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
2025
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
2026
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
2026
iGRPO: Self-Feedback-Driven LLM Reasoning
2026
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
2026
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
2026
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
2025
BroRL: Scaling Reinforcement Learning via Broadened Exploration
2025
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
2025
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
2025
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
2025
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
2023
Tailoring Self-Rationalizers with Multi-Reward Distillation
2023
Topics
cs.AI
cs.LG
cs.CL
RLHF & Alignment
Exploration
Model-Based RL
Policy Gradient
Fine-Tuning
Training Techniques
Evaluation