Abstract:This paper studies the optimization of Markov decision processes (MDPs) from a risk-seeking perspective, where the risk is measured by conditional value-at-risk (CVaR). The objective is to find a policy that maximizes the long-run CVaR of instantaneous rewards over an infinite horizon across all history-dependent randomized policies. By establishing two optimality inequalities of opposing directions, we prove that the maximum of long-run CVaR of MDPs over the set of history-dependent randomized policies can be found within the class of stationary randomized policies. In contrast to classical MDPs, we find that there may not exist an optimal stationary deterministic policy for maximizing CVaR. Instead, we prove the existence of an optimal stationary randomized policy that requires randomizing over at most two actions. Via a convex optimization representation of CVaR, we convert the long-run CVaR maximization MDP into a minimax problem, where we prove the interchangeability of minimum and maximum and the related existence of saddle point solutions. Furthermore, we propose an algorithm that finds the saddle point solution by solving two linear programs. These results are then extended to objectives that involve maximizing some combination of mean and CVaR of rewards simultaneously. Finally, we conduct numerical experiments to demonstrate the main results.

Average Optimality in Markov Decision Processes with Unbounded Rewards

On Average Optimality for Non-Stationary Markov Decision Processes in Borel Spaces

The Finiteness of the Reward Function and the Optimal Value Function in Markov Decision Processes

Continuous Time Markov Decision Processes with Expected Discounted Total Rewards

Average-Cost MDPs with Infinite State and Action Sets: New Sufficient Conditions for Optimality Inequalities and Equations

Continuous Time Markov Decision Processes with Nonuniformly Bounded Transition Rate: Expected Total Rewards

Analysis for Some Properties of Discrete Time Markov Decision Processes

Risk-sensitive discounted Markov decision processes with unbounded reward functions and Borel spaces

Risk-Sensitive Average Markov Decision Processes in General Spaces

Beyond Average Return in Markov Decision Processes

Finding Optimal Memoryless Policies of POMDPs under the Expected Average Reward Criterion

Beyond discounted returns: Robust Markov decision processes with average and Blackwell optimality

Optimal Stationary Policies for a Class of Countable Markov Control Processes

Solution to the risk-sensitive average cost optimality equation in a class of Markov decision processes with finite state space

On Linear Programming for Constrained and Unconstrained Average-Cost Markov Decision Processes with Countable Action Spaces and Strictly Unbounded Costs

A survey of recent results on continuous-time Markov decision processes

Optimal Sample Complexity for Average Reward Markov Decision Processes

OPTIMALITY STRATEGY OF AVERAGE COST BASED PERFORMANCE POTENTIALS FOR MARKOV CONTROL PROCESS

Finding Optimal Observation-Based Policies for Constrained POMDPs under the Expected Average Reward Criterion

On the Maximization of Long-Run Reward CVaR for Markov Decision Processes

On the optimality equation for average cost Markov decision processes and its validity for inventory control