A Novel Trajectory Planning Method Based on Trust Region Policy Optimization

Bao-Lin Ye,Jiajie Zhang,Lingxi Li,Weimin Wu
DOI: https://doi.org/10.1109/tiv.2024.3407484
IF: 8.2
2024-01-01
IEEE Transactions on Intelligent Vehicles
Abstract:Trajectory planning method is a research hotspot in autonomous driving. Existing reinforcement learning-based trajectory planning methods suffer from unstable performance due to the strong randomness of network weight parameter updates during the training process. Therefore, this paper proposes a novel trajectory planning method based on deep reinforcement learning trust region policy optimization (TRPO). Firstly, in order to enhance the robustness of the trajectory planning method based on deep reinforcement learning TRPO, a TRPO-LSTM based decision model was proposed. More specifically, a long short term memory (LSTM) based state feature extraction network was designed and embeded into a TRPO-based decision model to enhance the ability of TRPO to extract information from the environmental state space. Secondly, in order to make the planned trajectory adaptive to the dynamic changes of traffic environment, we presented a novel TRPO-LSTM trajectory fitting algorithm. To the best of our knowledge, this is the first work aiming at applying the TRPO-LSTM based decision model in the trajectory fitting process to search the optimal longitudinal trajectory speed. Finally, the proposed trajectory planning method was implemented and simulated on the CARLA simulator. The experimental results show that, compared with existing trajectory planning methods based on deep reinforcement learning algorithms, our proposed method achieves a cumulative reward improvement of over 28.9% in the scenario of four lane highway, and has better robustness. Meanwhile, the proposed method can achieve a lower collision rate of 0.93% while improving the average speed and comfort of vehicle driving.
What problem does this paper attempt to address?