TY - GEN
T1 - Attention-Enhanced Proximal Policy Optimization for Autonomous Robot Path Planning in Dynamic Environments
AU - Dai, Tingzhang
AU - Gu, Qingyu
AU - Kai, Xiayun
AU - Shi, Jiarong
AU - Lyu, Haoran
AU - Yao, Zhongyu
PY - 2026
Y1 - 2026
N2 - Autonomous robot path planning in dynamic, cluttered environments remains a significant problem in robotics. Traditional planning algorithms require exact prior knowledge and cannot effectively handle obstacle avoidance, while present-day deep learning (DL) techniques process all inputs uniformly, which is not suitable for efficient convergence. In this work, we propose AEPPO: Attention-Enhanced Proximal Policy Optimization for autonomous robots. AEPPO introduces a multi-head self-perception mechanism for observing spatial information with the help of attention heads, enabling the policy to focus more on dangerous object-related information. A reward function based on potential-field guidance and collision penalty is also easier to learn. In the MiniGrid and Gazebo simulated environments (static, corridor, dynamic-obstacle), the experimental results show that the success rate of AEPPO is 91.8% on MiniGrid with a 3.2% collision rate, and on the most difficult dynamic map, the success rate reaches 84.3%. Compared to five other baselines - A*, RRT, DQN, standard PPO, and Attention-DQN - training convergence results show that AEPPO reaches the target performance threshold approximately 27% faster than standalone standard PPO with much lower variance, demonstrating the effectiveness of attention-guided perception for safe and efficient robot navigation. © 2026 IEEE.
AB - Autonomous robot path planning in dynamic, cluttered environments remains a significant problem in robotics. Traditional planning algorithms require exact prior knowledge and cannot effectively handle obstacle avoidance, while present-day deep learning (DL) techniques process all inputs uniformly, which is not suitable for efficient convergence. In this work, we propose AEPPO: Attention-Enhanced Proximal Policy Optimization for autonomous robots. AEPPO introduces a multi-head self-perception mechanism for observing spatial information with the help of attention heads, enabling the policy to focus more on dangerous object-related information. A reward function based on potential-field guidance and collision penalty is also easier to learn. In the MiniGrid and Gazebo simulated environments (static, corridor, dynamic-obstacle), the experimental results show that the success rate of AEPPO is 91.8% on MiniGrid with a 3.2% collision rate, and on the most difficult dynamic map, the success rate reaches 84.3%. Compared to five other baselines - A*, RRT, DQN, standard PPO, and Attention-DQN - training convergence results show that AEPPO reaches the target performance threshold approximately 27% faster than standalone standard PPO with much lower variance, demonstrating the effectiveness of attention-guided perception for safe and efficient robot navigation. © 2026 IEEE.
KW - autonomous navigation
KW - deep reinforcement learning
KW - proximal policy optimization
KW - reward shaping
KW - robot path planning
KW - self-attention mechanism
UR - https://www.scopus.com/pages/publications/105044578273
UR - https://www.scopus.com/pages/publications/105044578273#tab=impact
U2 - 10.1109/ISCTIS70043.2026.11572138
DO - 10.1109/ISCTIS70043.2026.11572138
M3 - RGC 32 - Refereed conference paper (with host publication)
T3 - International Symposium on Computer Technology and Information Science, ISCTIS
SP - 1612
EP - 1616
BT - 2026 6th International Symposium on Computer Technology and Information Science (ISCTIS)
PB - IEEE
T2 - 2026 6th International Symposium on Computer Technology and Information Science, ISCTIS 2026
Y2 - 15 May 2026 through 17 May 2026
ER -