Intrinsically Motivated Hierarchical Policy Learning In Multi-objective Markov Decision Processes
2023 Β· Sherif Abdelfattah, Kathryn Merrick, Jiankun Hu
Abstract
Multi-objective Markov decision processes are sequential decision-making problems that involve multiple conflicting reward functions that cannot be optimized simultaneously without a compromise. This type of problems cannot be solved by a single optimal policy as in the conventional case. Alternatively, multi-objective reinforcement learning methods evolve a coverage set of optimal policies that can satisfy all possible preferences in solving the problem. However, many of these methods cannot generalize their coverage sets to work in non-stationary environments. In these environments, the parameters of the state transition and reward distribution vary over time. This limitation results in significant performance degradation for the evolved policy sets. In order to overcome this limitation, there is a need to learn a generic skill set that can bootstrap the evolution of the policy coverage set for each shift in the environment dynamics therefore, it can facilitate a continuous learning
Authors
(none)
Tags
Stats
Related papers
- Globally Optimal Hierarchical Reinforcement Learning For Linearly-solvable Markov Decision Processes (2021)2.26
- Multi-timescale Ensemble Q-learning For Markov Decision Process Policy Optimization (2024)6.34
- Policy Dispersion In Non-markovian Environment (2023)0.00
- Robust Batch Policy Learning In Markov Decision Processes (2020)0.00
- Sample-efficient Multi-objective Learning Via Generalized Policy Improvement Prioritization (2023)5.24
- Hierarchy Through Composition With Linearly Solvable Markov Decision Processes (2016)0.00
- Configurable Markov Decision Processes (2018)0.00
- Learning To Switch Among Agents In A Team Via 2-layer Markov Decision Processes (2020)0.00