This paper establishes that an MDP with a unique optimal policy and ergodic
associated transition matrix ensures the convergence of various versions of the
Value Iteration algorithm at a geometric rate that exceeds the discount factor
{\gamma} for both discounted and average-reward criteria.
Related papers
Ranked by semantic similarity β how closely each paper's abstract matches this one (100% = near-identical topic).