Abstract:Dynamic programming (DP) is an approach to computing the optimal control policy over time under nonlinearity and uncertainty by employing the principle of optimality introduced by Richard Bellman. Instead of enumerating all possible control sequences, dynamic programming only searches admissible state and/or action values that satisfy the principle of optimality. Therefore, the computation complexity can be much improved over the direct enumeration method. However, the computational efforts and the data storage requirement increase exponentially with the dimensionality of the system, which are reflected in the three curses: the state space, the observation space, and the action space. Thus, the traditional DP approach was limited to solving small size problems. This paper aims at providing an overview of latest development of a class of approximate/adaptive dynamic programming algorithms including those applicable to continuous state and continuous control problems. The paper will especially review direct heuristic dynamic programming (direct (HDP), its design and applications, which include large and complex continuous state and control problems. In addition to the basic principle of direct HDP, the paper includes two application studies of the direct HDP - one is when it is used in a nonlinear tracking problem, and the other is on a power grid coordination control problem based on China southern network.

Dynamic Programming: From Local Optimality to Global Optimality

The Principle of Optimality in Dynamic Programming: A Pedagogical Note

A Convex Optimization Approach to Dynamic Programming in Continuous State and Action Spaces

Optimal Control of Hybrid Systems Based on Dynamic Programming

Convergence of Dynamic Programming Models

Exact Dynamic Programming for Positive Systems with Linear Optimal Cost

Model-free Adaptive Dynamic Programming for Optimal Control of Discrete-time Affine Nonlinear System

Dynamic Programming: Finite States

Gradient-Bounded Dynamic Programming with Submodular and Concave Extensible Value Functions

The landscape of deterministic and stochastic optimal control problems: One-shot Optimization versus Dynamic Programming

Optimizing Operations Management and Business Analytics Strategies under Uncertainty: Dynamic Programming

Positive Dynamic Programming: A Critique

Dynamic Programming in Probability Spaces via Optimal Transport

On the Performance Bounds of some Policy Search Dynamic Programming Algorithms

Convergence Guarantee of Dynamic Programming for LTL Surrogate Reward

Dynamic Programming Deconstructed: Transformations of the Bellman Equation and Computational Efficiency

Online Abstract Dynamic Programming with Contractive Models

Stochastic Dynamic Programming with Non-linear Discounting

Approximate dynamic programming for continuous state and control problems

A Course in Dynamic Optimization