Abstract
Path planning is one of the core technologies enabling intelligent decision-making in autonomous navigation for mobile robots. However, existing reinforcement learning methods still face multiple challenges in complex environments, such as sparse rewards, insufficient exploration, slow convergence, and limited generalization capabilities. To address these issues, this paper proposes an improved version of the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm. The algorithm incorporates an n-step return mechanism to enhance the modeling of long-term dependencies, and introduces a dynamic hybrid exploration strategy that combines Gaussian noise with Ornstein–Uhlenbeck (OU) noise to improve exploration stability and training efficiency. In addition, a shaped reward function is constructed by integrating several key factors, including distance to the target, heading alignment, obstacle avoidance, and time efficiency, to provide continuous feedback during the learning process. The proposed method is validated on the Gazebo simulation platform. Experimental simulation results show that, compared to the original algorithm, the improved version increases the success rate by 37% in complex static environments and by 28% in dynamic obstacle scenarios. The approach significantly accelerates training convergence and enhances policy generalization, effectively demonstrating its practicality for robotic path planning tasks.
Author supplied keywords
Cite
CITATION STYLE
Yang, X., Wang, Q., Li, J., & Jiang, X. (2025). NM-TD3: A Hybrid Noise-Driven TD3 Algorithm With Long-Term Reward Propagation for Mobile Robot Path Planning. IEEE Access, 13, 149921–149932. https://doi.org/10.1109/ACCESS.2025.3602286
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.