Offline Reinforcement Learning With OOD State Correction And OOD Action Suppression
2024 Β· Yixiu Mao, Qi Wang, Chen Chen, et al.
Abstract
In offline reinforcement learning (RL), addressing the out-of-distribution (OOD) action issue has been a focus, but we argue that there exists an OOD state issue that also impairs performance yet has been underexplored. Such an issue describes the scenario when the agent encounters states out of the offline dataset during the test phase, leading to uncontrolled behavior and performance degradation. To this end, we propose SCAS, a simple yet effective approach that unifies OOD state correction and OOD action suppression in offline RL. Technically, SCAS achieves value-aware OOD state correction, capable of correcting the agent from OOD states to high-value in-distribution states. Theoretical and empirical results show that SCAS also exhibits the effect of suppressing OOD actions. On standard offline RL benchmarks, SCAS achieves excellent performance without additional hyperparameter tuning. Moreover, benefiting from its OOD state correction feature, SCAS demonstrates enhanced robustness
Authors
(none)
Tags
Stats
Related papers
- Beyond OOD State Actions: Supported Cross-domain Offline Reinforcement Learning (2023)0.00
- Alberdice: Addressing Out-of-distribution Joint Actions In Offline Multi-agent RL Via Alternating Stationary Distribution Correction Estimation (2023)0.00
- State-constrained Offline Reinforcement Learning (2024)0.00
- SAMG: Offline-to-online Reinforcement Learning Via State-action-conditional Offline Model Guidance (2024)0.00
- Mildly Conservative Q-learning For Offline Reinforcement Learning (2022)0.00
- Rethinking Out-of-distribution Detection For Reinforcement Learning: Advancing Methods For Evaluation And Detection (2024)2.26
- BRAC+: Improved Behavior Regularized Actor Critic For Offline Reinforcement Learning (2021)0.00
- Adaptive Advantage-guided Policy Regularization For Offline Reinforcement Learning (2024)3.09