Abstract:We study online policy optimization in nonlinear time-varying dynamical systems where the true dynamical models are unknown to the controller. This problem is challenging because, unlike in linear systems, the controller cannot obtain globally accurate estimations of the ground-truth dynamics using local exploration. We propose a meta-framework that combines a general online policy optimization algorithm ($\texttt{ALG}$) with a general online estimator of the dynamical system's model parameters ($\texttt{EST}$). We show that if the hypothetical joint dynamics induced by $\texttt{ALG}$ with known parameters satisfies several desired properties, the joint dynamics under inexact parameters from $\texttt{EST}$ will be robust to errors. Importantly, the final policy regret only depends on $\texttt{EST}$'s predictions on the visited trajectory, which relaxes a bottleneck on identifying the true parameters globally. To demonstrate our framework, we develop a computationally efficient variant of Gradient-based Adaptive Policy Selection, called Memoryless GAPS (M-GAPS), and use it to instantiate $\texttt{ALG}$. Combining M-GAPS with online gradient descent to instantiate $\texttt{EST}$ yields (to our knowledge) the first local regret bound for online policy optimization in nonlinear time-varying systems with unknown dynamics.

Online Adaptive Optimization Algorithm for Semi-Markov Control Processes

A nonlinear predictive control algorithm based on fuzzy online modeling and discrete optimization

Online policy iteration algorithm for semi-Markov switching state-space control processes

Optimization Of Semi-Markov Switching State-Space Control Processes For Network Communication Systems

Online Adaptation of Kernel Learning Adaptive Predictive Controller

Optimization Algorithms for Semi-Markov Control Processes with Average Criteria

Model-free Adaptive Dynamic Programming for Optimal Control of Discrete-time Affine Nonlinear System

Adaptive Learning-Based Path-Tracking Control for Unknown Vehicle Systems under Performance Optimization

Online Adaptive Optimal Control for Continuous-Time Nonlinear Systems with Completely Unknown Dynamics.

Adaptive Optimization of Time-out Policy for Dynamic Power Management Based on SMCP

Fuzzy-Based Adaptive Optimization of Unknown Discrete-Time Nonlinear Markov Jump Systems With Off-Policy Reinforcement Learning

Online Adaptive Optimal Control Algorithm Based on Synchronous Integral Reinforcement Learning With Explorations

Logic Switching Based Online Periodic Adaptive Learning Control Algorithm Dealing with Unknown Period and Bound of the Uncertain Parameter

Online Policy Optimization in Unknown Nonlinear Systems

Near Optimal Control for a Class of Stochastic Hybrid Systems.

Optimal control for continuous-time Markov jump singularly perturbed systems : A hybrid reinforcement learning scheme

An Online Algorithm for Solving Feedback Optimal Control Problems with Partial Observations

Online Reinforcement Learning-based Neural Network Controller Design for Affine Nonlinear Discrete-time Systems.

Data-driven cooperative optimal output regulation for linear discrete-time multi-agent systems by online distributed adaptive internal model approach

Model-Free Adaptive Optimal Control for Unknown Nonlinear Multiplayer Nonzero-Sum Game

Reinforcement Learning-Based Direct Adaptive Optimal Control of JLQ Model