Abstract:Reinforcement learning (RL) has shown a promising performance in learning optimal policies for a variety of sequential decision-making tasks. However, in many real-world RL problems, besides optimizing the main objectives, the agent is expected to satisfy a certain level of safety (e.g., avoiding collisions in autonomous driving). While RL problems are commonly formalized as Markov decision processes (MDPs), safety constraints are incorporated via constrained Markov decision processes (CMDPs). Although recent advances in safe RL have enabled learning safe policies in CMDPs, these safety requirements should be satisfied during both training and in the deployment process. Furthermore, it is shown that in memory-based and partially observable environments, these methods fail to maintain safety over unseen out-of-distribution observations. To address these limitations, we propose a Lyapunov-based uncertainty-aware safe RL model. The introduced model adopts a Lyapunov function that converts trajectory-based constraints to a set of local linear constraints. Furthermore, to ensure the safety of the agent in highly uncertain environments, an uncertainty quantification method is developed that enables identifying risk-averse actions through estimating the probability of constraint violations. Moreover, a Transformers model is integrated to provide the agent with memory to process long time horizons of information via the self-attention mechanism. The proposed model is evaluated in grid-world navigation tasks where safety is defined as avoiding static and dynamic obstacles in fully and partially observable environments. The results of these experiments show a significant improvement in the performance of the agent both in achieving optimality and satisfying safety constraints.

Curious iLQR: Resolving Uncertainty in Model-based RL

Model-Based Robot Learning Control with Uncertainty Directed Exploration

Curiosity model policy optimization for robotic manipulator tracking control with input saturation in uncertain environment

Curiosity-Driven Reinforcement Learning based Low-Level Flight Control

Curiosity driven reinforcement learning for motion planning on humanoids

Curious Meta-Controller: Adaptive Alternation between Model-Based and Model-Free Control in Deep Reinforcement Learning

Convergent iLQR for Safe Trajectory Planning and Control of Legged Robots

Lyapunov-based uncertainty-aware safe reinforcement learning

Disturbance Observer-based Control Barrier Functions with Residual Model Learning for Safe Reinforcement Learning

Model-free design of stochastic LQR controller from a primal–dual optimization perspective

Model-Based Reinforcement Learning Inspired by Augmented PD for Robotic Control

Reinforcement Learning for Safety-Critical Control under Model Uncertainty, using Control Lyapunov Functions and Control Barrier Functions

Uncertainty in Bayesian Reinforcement Learning for Robot Manipulation Tasks with Sparse Rewards.

Neural-iLQR: A Learning-Aided Shooting Method for Trajectory Optimization

i2LQR: Iterative LQR for Iterative Tasks in Dynamic Environments

Model-Free Design of Stochastic LQR Controller from Reinforcement Learning and Primal-Dual Optimization Perspective

Supervised Meta-Reinforcement Learning with Trajectory Optimization for Manipulation Tasks

Integrating DeepRL with Robust Low-Level Control in Robotic Manipulators for Non-Repetitive Reaching Tasks

Safe Deep Model-Based Reinforcement Learning with Lyapunov Functions

RL + Model-based Control: Using On-demand Optimal Control to Learn Versatile Legged Locomotion

Dynamic Neural Curiosity Enhances Learning Flexibility for Autonomous Goal Discovery