基于改进近端策略优化的航标自航方法研究

    Autonomous navigation method for aids to navigation based on improved proximal policy optimization

    • 摘要: 在近海港口航道繁忙水域,海上实体航标极易发生碰撞,导致标体损坏,无法正常运行。为解决现有方案针对航标自航研究领域空白,不能实现自主航行避碰,未能有效结合人工智能技术使其智能化航行工作等问题,提出一种改进近端策略优化并结合速度障碍的航标自航方法。近端策略优化算法采用离线学习方式提高网络更新效率,设计航标自航动作空间,能更快实现模型收敛;但在模型收敛时容易陷入局部最优,为避免模型陷入局部最优,结合速度障碍算法表示状态空间,用速度向量设计奖励函数,使其在计算距离奖励函数基础上,增加基于位置和速度信息的奖励函数,提高收敛速度,并满足航标在近海港口航道中自适应地选择最佳速度;为提高算法鲁棒性,增加熵奖励函数对近端策略优化算法改进,以激励更多探索。试验表明,增加熵奖励函数与未增加相比,在平均航行步长、平均航行速度等指标上分别提升了9.3%、10.2%,且模型测试成功率保持在99%以上。该方法实现了航标无人操控、自主航行及避碰的安全作业模式,降低航标在繁忙水域被碰撞损坏的风险,减少人工巡检和维护成本。对加快构建现代化航海保障体系、保障海上航行安全具有重要战略意义。

       

      Abstract: In the busy waters of nearshore port channels,physical Aids to Navigation(AtoN) are highly susceptible to collisions,which may cause structural damage and prevent them from operating normally. To address the research gap in autonomous navigation of AtoN,including the inability of existing methods to achieve autonomous navigation and collision avoidance and their insufficient integration with artificial intelligence technologies for intelligent navigation,this paper proposes an autonomous navigation method for AtoN based on improved Proximal Policy Optimization(PPO) combined with the Velocity Obstacle(VO) algorithm. The PPO algorithm was improved by introducing an offline learning strategy to enhance network update efficiency. By designing the autonomous navigation action space for AtoN,the model can achieve faster convergence. However,the model is prone to falling into local optima during convergence. To avoid this problem,the VO algorithm is incorporated to represent the state space,and a reward function is designed using velocity vectors. On the basis of calculating the distance-based reward function,an additional reward function based on position and velocity information is introduced to improve convergence speed and enable AtoN to adaptively select the optimal velocity in nearshore port channels. To further enhance algorithm robustness,an entropy reward term is added to the PPO algorithm to encourage greater exploration. Experimental results show that,compared with the method without the entropy reward term,the proposed method improves the average navigation step length and average navigation velocity by 9.3% and 10.2%,respectively,while maintaining a model test success rate of over 99%. The proposed method enables a safe operational mode featuring unmanned control,autonomous navigation,and collision avoidance for AtoN. It can reduce the collision damage risk of AtoN in busy waters and lower the costs of manual inspection and maintenance. The study is of significance for accelerating the construction of a modern maritime navigation support system and ensuring maritime navigation safety.

       

    /

    返回文章
    返回