Abstract:This paper investigates the optimization problem of an infinite stage discrete time Markov decision process (MDP) with a long‐run average metric considering both mean and variance of rewards together. Such performance metric is important since the mean indicates average returns and the variance indicates risk or fairness. However, the variance metric couples the rewards at all stages, the traditional dynamic programming is inapplicable as the principle of time consistency fails. We study this problem from a new perspective called the sensitivity‐based optimization theory. A performance difference formula is derived and it can quantify the difference of the mean‐variance combined metrics of MDPs under any two different policies. The difference formula can be utilized to generate new policies with strictly improved mean‐variance performance. A necessary condition of the optimal policy and the optimality of deterministic policies are derived. We further develop an iterative algorithm with a form of policy iteration, which is proved to converge to local optima both in the mixed and randomized policy space. Specially, when the mean reward is constant in policies, the algorithm is guaranteed to converge to the global optimum. Finally, we apply our approach to study the fluctuation reduction of wind power in an energy storage system, which demonstrates the potential applicability of our optimization method.

Mixed Markov Decision Processes in a Semi-Markov Environment with Discounted Criterion

Performance Optimization of Semi-Markov Decision Processes with Discounted-cost Criteria.

Continuous Time Markov Decision Processes with Discounted Moment Criterion

Continuous Time Markov Decision Processes with Expected Discounted Total Rewards

Performance Optimization for Countable Semi-Markov Decision Processes with Discounted-cost

Relations Between Discounted Models and Average Models for Semi-Markov Decision Processes

Optimal Replacement of a System According Toa Semi-Markov Decision Process in a Semi-Markov Environment

Towards Analysis Of Semi-Markov Decision Processes

Markov Decision Processes with Incomplete Information and Semi-Uniform Feller Transition Probabilities

Markov Decision Processes under Risk Sensitivity: A Discount Vanishing Approach

Analysis for Some Properties of Discrete Time Markov Decision Processes

Optimal Stationary Policies for Semi-Markov Control Processes with Discounted-Cost Criteria

Discrete Type Shock Semi-Markov Decision Processes with Borel State Space

Markov Decision Processes with Time-Varying Geometric Discounting

Optimal Policies for Quantum Markov Decision Processes

A survey of recent results on continuous-time Markov decision processes

Error bounds of optimization algorithms for semi-Markov decision processes

The Finiteness of the Reward Function and the Optimal Value Function in Markov Decision Processes

Continuous Time Markov Decision Processes with Nonuniformly Bounded Transition Rate: Expected Total Rewards

Risk-sensitive discounted Markov decision processes with unbounded reward functions and Borel spaces

Risk‐Sensitive Markov Decision Processes with Combined Metrics of Mean and Variance