In these notes we will tackle the problem of finding optimal policies for
Markov decision processes (MDPs) which are not fully known to us. Our intention
is to slowly transition from an offline setting to an online (learning)
setting. Namely, we are moving towards reinforcement learning.
Related papers
Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).