Inverse Reinforcement Learning In Contextual Mdps
2019 Β· Stav Belogolovsky, Philip Korsunsky, Shie Mannor, et al.
Abstract
We consider the task of Inverse Reinforcement Learning in Contextual Markov Decision Processes (MDPs). In this setting, contexts, which define the reward and transition kernel, are sampled from a distribution. In addition, although the reward is a function of the context, it is not provided to the agent. Instead, the agent observes demonstrations from an optimal policy. The goal is to learn the reward mapping, such that the agent will act optimally even when encountering previously unseen contexts, also known as zero-shot transfer. We formulate this problem as a non-differential convex optimization problem and propose a novel algorithm to compute its subgradients. Based on this scheme, we analyze several methods both theoretically, where we compare the sample complexity and scalability, and empirically. Most importantly, we show both theoretically and empirically that our algorithms perform zero-shot transfer (generalize to new and unseen contexts). Specifically, we present empirical e
Authors
(none)
Tags
Stats
Related papers
- No-regret Exploration In Contextual Reinforcement Learning (2019)0.00
- Reinforcement Learning In Reward-mixing Mdps (2021)0.00
- Learning Non-markovian Reward Models In Mdps (2020)0.00
- Reinforcement Learning In Presence Of Discrete Markovian Context Evolution (2022)0.00
- Markov Decision Processes With Continuous Side Information (2017)0.00
- Exploration Implies Data Augmentation: Reachability And Generalisation In Contextual Mdps (2024)0.00
- Policy Dispersion In Non-markovian Environment (2023)0.00
- Reward Is Enough For Convex Mdps (2021)0.00