Abstract:We study sequential decision-making problems in which each agent aims to maximize the expected total reward while satisfying a constraint on the expected total utility. We employ the natural policy gradient method to solve the discounted inﬁnite-horizon Constrained Markov Decision Processes (CMDPs) problem. Speciﬁcally, we propose a new Natural Policy Gradient Primal-Dual (NPG-PD) method for CMDPs which updates the primal variable via natural policy gradient ascent and the dual variable via projected sub-gradient descent. Even though the underlying maximization involves a nonconcave objective function and a nonconvex constraint set under the softmax policy parametrization, we prove that our method achieves global convergence with sublinear rates regarding both the optimality gap and the constraint violation. Such a convergence is independent of the size of the state-action space, i.e., it is dimension-free. Furthermore, for the general smooth policy class, we establish sublinear rates of convergence regarding both the optimality gap and the constraint violation, up to a function approximation error caused by restricted policy parametrization. Finally, we show that two sample-based NPG-PD algorithms inherit such non-asymptotic convergence properties and provide ﬁnite-sample complexity guarantees. To the best of our knowledge, our work is the ﬁrst to establish non-asymptotic convergence guarantees of policy-based primal-dual methods for solving inﬁnite-horizon discounted CMDPs. We also provide computational results to demonstrate merits of our approach.

A Simulation Optimization Algorithm for CTMDPs Based on Randomized Stationary Policies

Simulation Optimization Algorithm for SMDPs with Parameterized Randomized Stationary Policies

Stochastic Steepest-Descent Optimization Of Multiple-Objective Mobile Sensor Coverage

Parallel Optimization for Markov Control Processes Based on Performance Potentials Simulation

Optimal Time-Abstract Schedulers for CTMDPs and Markov Games

Optimal Policies for a Continuous Time MCP with Compact Action Set

Performance Optimization for Countable Semi-Markov Decision Processes with Discounted-cost

Performance Optimization of Continuous-Time Markov Control Processes Based on Performance Potentials

Error bounds of optimization algorithms for semi-Markov decision processes

Potential Based Optimization Algorithm Of Constrained Markov Decision Processes

Performance Potentials Based Stochastic Optimization and Parallel Algorithm for a Class of CQN

Two-Timescale Simulation-based Algorithm for Markov Decision Process Based on Performance Potentials

Optimal Stationary Policies for Semi-Markov Control Processes with Discounted-Cost Criteria

The policy gradient estimation of continuous-time hidden Markov decision processes

Optimization Algorithms for Semi-Markov Control Processes with Average Criteria

Performance Optimization of Semi-Markov Decision Processes with Discounted-cost Criteria.

Event-based optimization for finite-horizon total-cost markov decision processes

A safe exploration approach to constrained Markov decision processes

Continuous-time Markov Decision Processes to Call Admission Control Problem

Simulation Optimization: A New Approach

Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision Processes