The Effects Of Reward Misspecification: Mapping And Mitigating Misaligned Models
2022 Β· Alexander Pan, Kush Bhatia, Jacob Steinhardt
Abstract
Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied. To understand how reward hacking arises, we construct four RL environments with misspecified rewards. We investigate reward hacking as a function of agent capabilities: model capacity, action space resolution, observation space noise, and training time. More capable agents often exploit reward misspecifications, achieving higher proxy reward and lower true reward than less capable agents. Moreover, we find instances of phase transitions: capability thresholds at which the agent's behavior qualitatively shifts, leading to a sharp decrease in the true reward. Such phase transitions pose challenges to monitoring the safety of ML systems. To address this, we propose an anomaly detection task for aberrant policies and offer several baseline detectors.
Authors
(none)
Tags
Stats
Related papers
- Correlated Proxies: A New Definition And Improved Mitigation For Reward Hacking (2024)2.76
- Reward Hacking Benchmark: Measuring Exploits In LLM Agents With Tool Use (2026)0.00
- Causal Confusion And Reward Misidentification In Preference-based Reward Learning (2022)0.00
- REBEL: Reward Regularization-based Approach For Robotic Reinforcement Learning From Human Feedback (2023)0.00
- Disturbing Reinforcement Learning Agents With Corrupted Rewards (2021)0.00
- On The Model-misspecification In Reinforcement Learning (2023)0.00
- Reward Tampering Problems And Solutions In Reinforcement Learning: A Causal Influence Diagram Perspective (2019)0.00
- Learning The Reward Function For A Misspecified Model (2018)0.00