Abstract:Reinforcement learning (RL) has shown a promising performance in learning optimal policies for a variety of sequential decision-making tasks. However, in many real-world RL problems, besides optimizing the main objectives, the agent is expected to satisfy a certain level of safety (e.g., avoiding collisions in autonomous driving). While RL problems are commonly formalized as Markov decision processes (MDPs), safety constraints are incorporated via constrained Markov decision processes (CMDPs). Although recent advances in safe RL have enabled learning safe policies in CMDPs, these safety requirements should be satisfied during both training and in the deployment process. Furthermore, it is shown that in memory-based and partially observable environments, these methods fail to maintain safety over unseen out-of-distribution observations. To address these limitations, we propose a Lyapunov-based uncertainty-aware safe RL model. The introduced model adopts a Lyapunov function that converts trajectory-based constraints to a set of local linear constraints. Furthermore, to ensure the safety of the agent in highly uncertain environments, an uncertainty quantification method is developed that enables identifying risk-averse actions through estimating the probability of constraint violations. Moreover, a Transformers model is integrated to provide the agent with memory to process long time horizons of information via the self-attention mechanism. The proposed model is evaluated in grid-world navigation tasks where safety is defined as avoiding static and dynamic obstacles in fully and partially observable environments. The results of these experiments show a significant improvement in the performance of the agent both in achieving optimality and satisfying safety constraints.

Constrained Cross-Entropy Method for Safe Reinforcement Learning

Safe Reinforcement Learning Using Finite-Horizon Gradient-based Estimation

Successive Convex Approximation Based Off-Policy Optimization for Constrained Reinforcement Learning

Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints

Train Trajectory Optimization with High-Risk State Space Boundaries: A Safe Reinforcement Learning Approach

Safe Reinforcement Learning via Hierarchical Adaptive Chance-Constraint Safeguards

Reachability Constrained Reinforcement Learning.

Probabilistic Constraint for Safety-Critical Reinforcement Learning

Constraints Penalized Q-learning for Safe Offline Reinforcement Learning.

Long and Short-Term Constraints Driven Safe Reinforcement Learning for Autonomous Driving

Lyapunov-based uncertainty-aware safe reinforcement learning

Model-Based Chance-Constrained Reinforcement Learning via Separated Proportional-Integral Lagrangian

Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism

Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic Environments

On Bellman's principle of optimality and Reinforcement learning for safety-constrained Markov decision process

Model-Based Actor-Critic with Chance Constraint for Stochastic System

Safe Model-Based Reinforcement Learning with an Uncertainty-Aware Reachability Certificate

Model-based Safe Deep Reinforcement Learning via a Constrained Proximal Policy Optimization Algorithm

Towards Safe Reinforcement Learning Via Constraining Conditional Value-at-Risk

Learn Zero-Constraint-Violation Policy in Model-Free Constrained Reinforcement Learning

Revisiting Safe Exploration in Safe Reinforcement learning