Awesome Large Language Models
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
David Krueger — most-cited papers & profile · Large Language Models
← authors
·
overview
David Krueger
18
papers ·
273
citations ·
26
h-index
University of Cambridge · Indian Institute of Technology Hyderabad
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Scalable agent alignment via reward modeling: a research direction
2018 · 126 citations
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023 · 93 citations
Goal Misgeneralization in Deep Reinforcement Learning
2021 · 23 citations
Uncertainty in Multitask Transfer Learning
2018 · 12 citations
Active Reinforcement Learning: Observing Rewards at a Cost
2020 · 12 citations
Reward Model Ensembles Help Mitigate Overoptimization
2023 · 2 citations
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
2024 · 2 citations
Predicting Future Actions of Reinforcement Learning Agents
2024 · 2 citations
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
2024 · 1 citations
Mitigating Goal Misgeneralization via Minimax Regret
2025
Interpreting Emergent Planning in Model-Free Reinforcement Learning
2025
Deep Prior
2017
Neural Autoregressive Flows
2018
Domain Generalization for Robust Model-Based Offline Reinforcement Learning
2022
Thinker: Learning to Plan and Act
2023
Top co-authors
Brandon Jaipersaud
· 1
Dmitrii Krasheninnikov
· 1
Ekdeep Singh Lubana
· 1
Joe Kwon
· 1
Madeline Brumley
· 1
Usman Anwar
· 1
Topics
Model-Based RL
RLHF & Alignment
Safe RL
Value-Based
Game AI
Exploration
stat.ML
Meta-RL
Policy Gradient
Training & Sampling