Abstract:Ensuring safety of nonlinear systems under model uncertainty and external disturbances is crucial, especially for real-world control tasks. Predictive methods such as robust model predictive control (RMPC) require solving nonconvex optimization problems online, which leads to high computational burden and poor scalability. Reinforcement learning (RL) works well with complex systems, but pays the price of losing rigorous safety guarantee. This paper presents a theoretical framework that bridges the advantages of both RMPC and RL to synthesize safety filters for nonlinear systems with state- and action-dependent uncertainty. We decompose the robust invariant set (RIS) into two parts: a target set that aligns with terminal region design of RMPC, and a reach-avoid set that accounts for the rest of RIS. We propose a policy iteration approach for robust reach-avoid problems and establish its monotone convergence. This method sets the stage for an adversarial actor-critic deep RL algorithm, which simultaneously synthesizes a reach-avoid policy network, a disturbance policy network, and a reach-avoid value network. The learned reach-avoid policy network is utilized to generate nominal trajectories for online verification, which filters potentially unsafe actions that may drive the system into unsafe regions when worst-case disturbances are applied. We formulate a second-order cone programming (SOCP) approach for online verification using system level synthesis, which optimizes for the worst-case reach-avoid value of any possible trajectories. The proposed safety filter requires much lower computational complexity than RMPC and still enjoys persistent robust safety guarantee. The effectiveness of our method is illustrated through a numerical example.

Q-learning-based Model-free Safety Filter

Learning Predictive Safety Filter via Decomposition of Robust Invariant Set

Learning Observation-Based Certifiable Safe Policy for Decentralized Multi-Robot Navigation

Safety Filtering While Training: Improving the Performance and Sample Efficiency of Reinforcement Learning Agents

Safe and Efficient Reinforcement Learning Using Disturbance-Observer-Based Control Barrier Functions

Data-Driven Safety Filter: An Input-Output Perspective

A Finite-Time Safety Filter for Learning-Based Autonomous Driving

Look Before You Leap: Safe Model-Based Reinforcement Learning with Human Intervention

Conformal Predictive Safety Filter for RL Controllers in Dynamic Environments

Safety Filters for Black-Box Dynamical Systems by Learning Discriminating Hyperplanes

Safe Reinforcement Learning via a Model-Free Safety Certifier

Model-Based Safe Reinforcement Learning with Time-Varying State and Control Constraints: An Application to Intelligent Vehicles

Safe Model-Based Reinforcement Learning for Systems with Parametric Uncertainties

Model-Based Safe Reinforcement Learning With Time-Varying Constraints: Applications to Intelligent Vehicles

Modular Safety-Critical Control of Legged Robots

Gameplay Filters: Robust Zero-Shot Safety through Adversarial Imagination

Safe Reinforcement Learning Filter for Multicopter Collision-Free Tracking under disturbances

Disturbance Observer-based Control Barrier Functions with Residual Model Learning for Safe Reinforcement Learning

Reinforcement Learning for Safe Robot Control using Control Lyapunov Barrier Functions

Model-free Neural Lyapunov Control for Safe Robot Navigation

Safe Reinforcement Learning Using Robust Control Barrier Functions