Abstract:Reinforcement learning (RL) agents are vulnerable to adversarial disturbances, which can deteriorate task performance or break down safety specifications. Existing methods either address safety requirements under the assumption of no adversary (e.g., safe RL) or only focus on robustness against performance adversaries (e.g., robust RL). Learning one policy that is both safe and robust under any adversaries remains a challenging open problem. The difficulty is how to tackle two intertwined aspects in the worst cases: feasibility and optimality. The optimality is only valid inside a feasible region (i.e., robust invariant set), while the identification of maximal feasible region must rely on how to learn the optimal policy. To address this issue, we propose a systematic framework to unify safe RL and robust RL, including the problem formulation, iteration scheme, convergence analysis and practical algorithm design. The unification is built upon constrained two-player zero-sum Markov games, in which the objective for protagonist is twofold. For states inside the maximal robust invariant set, the goal is to pursue rewards under the condition of guaranteed safety; for states outside the maximal robust invariant set, the goal is to reduce the extent of constraint violation. A dual policy iteration scheme is proposed, which simultaneously optimizes a task policy and a safety policy. We prove that the iteration scheme converges to the optimal task policy which maximizes the twofold objective in the worst cases, and the optimal safety policy which stays as far away from the safety boundary. The convergence of safety policy is established by exploiting the monotone contraction property of safety self-consistency operators, and that of task policy depends on the transformation of safety constraints into state-dependent action spaces. By adding two adversarial networks (one is for safety guarantee and the other is for task performance), we propose a practical deep RL algorithm for constrained zero-sum Markov games, called dually robust actor-critic (DRAC). The evaluations with safety-critical benchmarks demonstrate that DRAC achieves high performance and persistent safety under all scenarios (no adversary, safety adversary, performance adversary), outperforming all baselines by a large margin.

Progressive Adaptive Chance-Constrained Safeguards for Reinforcement Learning.

Safe Reinforcement Learning via Hierarchical Adaptive Chance-Constraint Safeguards

Cautious Adaptation For Reinforcement Learning in Safety-Critical Settings

Successive Convex Approximation Based Off-Policy Optimization for Constrained Reinforcement Learning

Model-Based Actor-Critic with Chance Constraint for Stochastic System

Safeguarded Progress in Reinforcement Learning: Safe Bayesian Exploration for Control Policy Synthesis

GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model

SAAC: Safe Reinforcement Learning as an Adversarial Game of Actor-Critics

Probabilistic Constraint for Safety-Critical Reinforcement Learning

Uniformly Safe RL with Objective Suppression for Multi-Constraint Safety-Critical Applications

SCPO: Safe Reinforcement Learning with Safety Critic Policy Optimization

Reinforcement Learning with Adaptive Regularization for Safe Control of Critical Systems

Safe Reinforcement Learning with Dual Robustness

Probabilistic Safeguard for Reinforcement Learning Using Safety Index Guided Gaussian Process Models

Model-Based Chance-Constrained Reinforcement Learning via Separated Proportional-Integral Lagrangian

Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic Environments

Long and Short-Term Constraints Driven Safe Reinforcement Learning for Autonomous Driving

Context-Aware Safe Reinforcement Learning for Non-Stationary Environments

ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning

Flipping-based Policy for Chance-Constrained Markov Decision Processes

Implicit Safe Set Algorithm for Provably Safe Reinforcement Learning