谢伟宁,陈龙胜,何国毅,等. 基于CLP-DDPG算法的复杂环境下无人机路径规划(特邀)[J]. 南昌航空大学学报(自然科学版),2026,40(2):1-10. doi: 10.3969/j.issn.2096-8566.2026.02.001
引用本文: 谢伟宁,陈龙胜,何国毅,等. 基于CLP-DDPG算法的复杂环境下无人机路径规划(特邀)[J]. 南昌航空大学学报(自然科学版),2026,40(2):1-10. doi: 10.3969/j.issn.2096-8566.2026.02.001
XIE Weining,CHEN Longsheng,HE Guoyi,et al. Unmanned aerial vehicle path planning based on CLP-DDPG algorithm in complex environments (invited)[J]. Journal of Nanchang Hangkong University (Natural Sciences),2026,40(2):1-10. doi: 10.3969/j.issn.2096-8566.2026.02.001
Citation: XIE Weining,CHEN Longsheng,HE Guoyi,et al. Unmanned aerial vehicle path planning based on CLP-DDPG algorithm in complex environments (invited)[J]. Journal of Nanchang Hangkong University (Natural Sciences),2026,40(2):1-10. doi: 10.3969/j.issn.2096-8566.2026.02.001

基于CLP-DDPG算法的复杂环境下无人机路径规划(特邀)

Unmanned Aerial Vehicle Path Planning Based on CLP-DDPG Algorithm in Complex Environments (Invited)

  • 摘要: 针对复杂环境下无人机路径规划存在的探索效率低、收敛速度慢与路径平滑度不佳等问题,本文提出一种融合课程学习与嵌入优先回放机制的深度确定性策略梯度算法的路径规划改进方法。首先,通过设计一套从易到难的障碍物环境课程,引导无人机逐步学习从简单到复杂的路径规划任务,提高训练效率和稳定性。然后,将优先回放机制嵌入算法中,保证在冗杂的经验池中快速提取有效经验,进一步提升收敛速度和稳定性,确保路径效率更高和平滑度更优。仿真结果表明,融合课程学习与嵌入优先回放机制的强化学习方法能有效提升无人机在未知复杂环境下的自主避障与路径规划能力,相比于传统的DDPG算法路径规划效率提高了 26.84\text\text% ,平滑度提升了 66.19\text\text% ,训练的收敛速度更快且后期训练的稳定性更好。

     

    Abstract: To address the issues of low exploration efficiency, slow convergence speed and poor path smoothness in unmanned aerial vehicle (UAV) path planning under complex environments, an improved path planning algorithm based on the deep deterministic policy gradient (DDPG) framework is proposed, which integrates curriculum learning and an embedded prioritized replay mechanism. First, a curriculum of obstacle environments ranging from easy to difficult is designed to guide the UAV to gradually learn path planning tasks from simple to complex, thereby improving training efficiency and stability. Then, the prioritized replay mechanism is embedded into the algorithm to ensure rapid extraction of effective experiences from the complex experience replay buffer, further enhancing convergence speed and stability, and ensuring higher path efficiency and better smoothness. Simulation results demonstrate that the reinforcement learning method integrating curriculum learning and the embedded prioritized replay mechanism can effectively improve the autonomous obstacle avoidance and path planning capabilities of UAVs in unknown complex environments. Compared with the traditional DDPG algorithm, the proposed method increases path planning efficiency by 26.84% and improves path smoothness by 66.19%, while achieving faster convergence speed during training and better stability in the later stage of training.

     

/

返回文章
返回