Abstract:In this work, we combined the model based reinforcement learning ( MBRL ) and model free reinforcement learning ( MFRL ) to stabilize a biped robot ( NAO robot ) on a rotating platform, where the angular velocity of the platform is unknown for the proposed learning algorithm and treated as the external disturbance. Nonparametric Gaussian processes normally require a large number of training data points to deal with the discontinuity of the estimated model. Although some improved method such as probabilistic inference for learning control ( PILCO ) does not require an explicit global model as the actions are obtained by directly searching the policy space, the overfitting and lack of model complexity may still result in a large deviation between the prediction and the real system. Besides, none of these approaches consider the data error and measurement noise during the training process and test process, respectively. We propose a hierarchical Gaussian processes ( GP ) models, containing two layers of independent GPs, where the physically continuous probability transition model of the robot is obtained. Due to the physically continuous estimation, the algorithm overcomes the overfitting problem with a guaranteed model complexity, and the number of training data is also reduced. The policy for any given initial state is generated automatically by minimizing the expected cost according to the predefined cost function and the obtained probability distribution of the state. Furthermore, a novel Q ( λ ) based MFRL method scheme is employed to improve the policy. Simulation results show that the proposed RL algorithm is able to balance NAO robot on a rotating platform, and it is capable of adapting to the platform with varying angular velocity.

A Q-learning approach to the continuous control problem of robot inverted pendulum balancing

Learning biped locomotion based on Q-learning and neural networks

Adaptive Control of an Inverted Pendulum by a Reinforcement Learning-based LQR Method

A Q-learning Method for Continuous Space Based on Self-organizing Fuzzy RBF Network

Vague Neural Network Based Reinforcement Learning Control System For Inverted Pendulum

Swing-up and Balance Control of Inverted Pendulum Based on Reinforcement Learning

Control of the Double Inverted Pendulum Based on Reinforcement Learning

Balance Controller Design for Inverted Pendulum Considering Detail Reward Function and Two-Phase Learning Protocol

Model-Based Reinforcement Learning In Continuous Environments Using Real-Time Constrained Optimization

Balance Control Of Robot With Cmac Based Q-Learning

Balance control of a biped robot on a rotating platform based on efficient reinforcement learning

A Kernel-Based Reinforcement Learning Approach To Stochastic Pole Balancing Control Systems

Q-Learning for Continuous State Space Based on a Support Vector Machine

Solve the Inverted Pendulum Problem Base on DQN Algorithm

A Hybrid Approach for Reinforcement Learning Using Virtual Policy Gradient for Balancing an Inverted Pendulum

Active Exploration Planning in Reinforcement Learning for Inverted Pendulum System Control

Learning Robust and Adaptive Real-World Continuous Control Using Simulation and Transfer Learning

Learning-Based Trajectory Tracking and Balance Control for Bicycle Robots with a Pendulum: A Gaussian Process Approach

Control of double linear inverted pendulum

Acquisition of A Gymnast-Like Robotic Giant-Swing Motion by Q-Learning and Improvement of the Repeatability

Limit cycles in inverted pendulum system by reinforcement learning