Abstract:Ensuring safety of nonlinear systems under model uncertainty and external disturbances is crucial, especially for real-world control tasks. Predictive methods such as robust model predictive control (RMPC) require solving nonconvex optimization problems online, which leads to high computational burden and poor scalability. Reinforcement learning (RL) works well with complex systems, but pays the price of losing rigorous safety guarantee. This paper presents a theoretical framework that bridges the advantages of both RMPC and RL to synthesize safety filters for nonlinear systems with state- and action-dependent uncertainty. We decompose the robust invariant set (RIS) into two parts: a target set that aligns with terminal region design of RMPC, and a reach-avoid set that accounts for the rest of RIS. We propose a policy iteration approach for robust reach-avoid problems and establish its monotone convergence. This method sets the stage for an adversarial actor-critic deep RL algorithm, which simultaneously synthesizes a reach-avoid policy network, a disturbance policy network, and a reach-avoid value network. The learned reach-avoid policy network is utilized to generate nominal trajectories for online verification, which filters potentially unsafe actions that may drive the system into unsafe regions when worst-case disturbances are applied. We formulate a second-order cone programming (SOCP) approach for online verification using system level synthesis, which optimizes for the worst-case reach-avoid value of any possible trajectories. The proposed safety filter requires much lower computational complexity than RMPC and still enjoys persistent robust safety guarantee. The effectiveness of our method is illustrated through a numerical example.

Synthesize Efficient Safety Certificates for Learning-Based Safe Control Using Magnitude Regularization

Learning-Based Safe Control for Robot and Autonomous Vehicle Using Efficient Safety Certificate

Joint Synthesis of Safety Certificate and Safe Control Policy Using Constrained Reinforcement Learning

Learning Observation-Based Certifiable Safe Policy for Decentralized Multi-Robot Navigation

Semi-Supervised Safe Visuomotor Policy Synthesis using Barrier Certificates

MAGICS: Adversarial RL with Minimax Actors Guided by Implicit Critic Stackelberg for Convergent Neural Synthesis of Robot Safety

Safe Control With Learned Certificates: A Survey of Neural Lyapunov, Barrier, and Contraction Methods for Robotics and Control

Safe Reinforcement Learning via a Model-Free Safety Certifier

Safe Nonlinear Control Using Robust Neural Lyapunov-Barrier Functions

Safe Online Dynamics Learning with Initially Unknown Models and Infeasible Safety Certificates

End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks

Learning Predictive Safety Filter via Decomposition of Robust Invariant Set

Scalable Synthesis of Verified Controllers in Deep Reinforcement Learning

Model-Free Safe Reinforcement Learning Through Neural Barrier Certificate

Synthesis and verification of robust-adaptive safe controllers

Zero-shot Transferable and Persistently Feasible Safe Control for High Dimensional Systems by Consistent Abstraction

Sablas: Learning Safe Control for Black-Box Dynamical Systems

Implicit Safe Set Algorithm for Provably Safe Reinforcement Learning

Learning Safe, Generalizable Perception-Based Hybrid Control With Certificates

Barrier Certified Safety Learning Control: When Sum-of-Square Programming Meets Reinforcement Learning