PAC Bounds For Imitation And Model-based Batch Learning Of Contextual Markov Decision Processes
2020 Β· Yash Nair, Finale Doshi-Velez
Abstract
We consider the problem of batch multi-task reinforcement learning with observed context descriptors, motivated by its application to personalized medical treatment. In particular, we study two general classes of learning algorithms: direct policy learning (DPL), an imitation-learning based approach which learns from expert trajectories, and model-based learning. First, we derive sample complexity bounds for DPL, and then show that model-based learning from expert actions can, even with a finite model class, be impossible. After relaxing the conditions under which the model-based approach is expected to learn by allowing for greater coverage of state-action space, we provide sample complexity bounds for model-based learning with finite model classes, showing that there exist model classes with sample complexity exponential in their statistical complexity. We then derive a sample complexity upper bound for model-based learning based on a measure of concentration of the data distribution
Authors
(none)
Tags
Stats
Related papers
- Model-based RL In Contextual Decision Processes: PAC Bounds And Exponential Improvements Over Model-free Approaches (2018)0.00
- Contextual Decision Processes With Low Bellman Rank Are Pac-learnable (2016)0.00
- Markov Decision Processes With Continuous Side Information (2017)0.00
- Robust Batch Policy Learning In Markov Decision Processes (2020)0.00
- Inverse Reinforcement Learning In Contextual Mdps (2019)8.82
- Reinforcement Learning In Presence Of Discrete Markovian Context Evolution (2022)0.00
- Sequence Model Imitation Learning With Unobserved Contexts (2022)0.00
- Batch Policy Learning In Average Reward Markov Decision Processes (2020)0.00