← all papers Β· overview

Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs

Abstract

This paper focuses on learning a Constrained Markov Decision Process (CMDP) via general parameterized policies. We propose a Primal-Dual based Regularized Accelerated Natural Policy Gradient (PDR-ANPG) algorithm that uses entropy and quadratic regularizers to reach this goal. For parameterized policy classes with a transferred compatibility approximation error, , PDR-ANPG achieves a last-iterate optimality gap and constraint violation with a sample complexity of . If the class is incomplete (), then the sample complexity reduces to for . Moreover, for complete policies with , our algorithm achieves a last-iterate optimality gap and constraint violation with sample complexity. It is a significant improvement over the state-of-the-art last-iterate guarantees of general parameterized CMDPs.

Related papers

Ranked by semantic similarity β€” how closely each paper's abstract matches this one (100% = near-identical topic).