Abstract:Reinforcement learning (RL) has shown a promising performance in learning optimal policies for a variety of sequential decision-making tasks. However, in many real-world RL problems, besides optimizing the main objectives, the agent is expected to satisfy a certain level of safety (e.g., avoiding collisions in autonomous driving). While RL problems are commonly formalized as Markov decision processes (MDPs), safety constraints are incorporated via constrained Markov decision processes (CMDPs). Although recent advances in safe RL have enabled learning safe policies in CMDPs, these safety requirements should be satisfied during both training and in the deployment process. Furthermore, it is shown that in memory-based and partially observable environments, these methods fail to maintain safety over unseen out-of-distribution observations. To address these limitations, we propose a Lyapunov-based uncertainty-aware safe RL model. The introduced model adopts a Lyapunov function that converts trajectory-based constraints to a set of local linear constraints. Furthermore, to ensure the safety of the agent in highly uncertain environments, an uncertainty quantification method is developed that enables identifying risk-averse actions through estimating the probability of constraint violations. Moreover, a Transformers model is integrated to provide the agent with memory to process long time horizons of information via the self-attention mechanism. The proposed model is evaluated in grid-world navigation tasks where safety is defined as avoiding static and dynamic obstacles in fully and partially observable environments. The results of these experiments show a significant improvement in the performance of the agent both in achieving optimality and satisfying safety constraints.

Robust Reinforcement Learning with UUB Guarantee for Safe Motion Control of Autonomous Robots

Learning Observation-Based Certifiable Safe Policy for Decentralized Multi-Robot Navigation

H_∞ Model-free Reinforcement Learning with Robust Stability Guarantee

Optimal Control for Constrained Discrete-Time Nonlinear Systems Based on Safe Reinforcement Learning.

Reinforcement Learning Control of Constrained Dynamic Systems with Uniformly Ultimate Boundedness Stability Guarantee

Robust Safe Reinforcement Learning under Adversarial Disturbances

Reinforcement Learning for Safe Robot Control using Control Lyapunov Barrier Functions

Safe Reinforcement Learning With Stability Guarantee for Motion Planning of Autonomous Vehicles

Model-Based Safe Reinforcement Learning with Time-Varying State and Control Constraints: An Application to Intelligent Vehicles

Model-Based Safe Reinforcement Learning With Time-Varying Constraints: Applications to Intelligent Vehicles

Navigation for Autonomous Vehicles Via Fast-Stable and Smooth Reinforcement Learning

Safe reinforcement learning for probabilistic reachability and safety specifications: A Lyapunov-based approach

Learning Robust Policies via Interpretable Hamilton-Jacobi Reachability-Guided Disturbances

Lyapunov-based uncertainty-aware safe reinforcement learning

Robust Reinforcement Learning for Risk-Sensitive Linear Quadratic Gaussian Control

Control of UAV Quadrotor Using Reinforcement Learning and Robust Controller

Safe Reinforcement Learning Using Robust Control Barrier Functions

End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks

Off-Policy Risk-Sensitive Reinforcement Learning-Based Constrained Robust Optimal Control

SSRL: A Safe and Smooth Reinforcement Learning Approach for Collision Avoidance in Navigation

Safe Model-Based Reinforcement Learning with an Uncertainty-Aware Reachability Certificate