{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T01:26:59Z","timestamp":1785893219145,"version":"3.56.0"},"reference-count":42,"publisher":"American Association for the Advancement of Science (AAAS)","content-domain":{"domain":["spj.science.org"],"crossmark-restriction":true},"short-container-title":["Intell Comput"],"published-print":{"date-parts":[[2025,1]]},"abstract":"<jats:p>In the field of path planning, the efficiency and effectiveness of deep reinforcement learning (DRL) methods are often constrained by the algorithms\u2019 exploration capabilities, particularly in dynamic and nondeterministic environments. This paper introduces a novel DRL optimization approach predicated on an action curiosity mechanism, designed to enhance both performance and efficiency in uncertain settings. By incentivizing agents to explore their surroundings more effectively, the action curiosity module amplifies learning efficiency and curtails training duration. The method\u2019s adaptability and stability in intricate or dynamic scenarios are augmented through a dynamically adjusted reward mechanism. To mitigate the issue of strategy degradation stemming from excessive exploration, we incorporate a cosine annealing strategy that finesses parameter adjustments in real time. Extensive experimentation reveals that our enhanced algorithm outperforms conventional methods markedly in terms of success rate and average reward, among other metrics. These experimental outcomes corroborate the proposed method\u2019s robustness and efficacy, laying a solid groundwork for efficient adaptive autonomous path planning within complex and nondeterministic environments.<\/jats:p>","DOI":"10.34133\/icomputing.0140","type":"journal-article","created":{"date-parts":[[2025,5,16]],"date-time":"2025-05-16T18:52:40Z","timestamp":1747421560000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark_01","source":"Crossref","is-referenced-by-count":16,"title":["Action-Curiosity-Based Deep Reinforcement Learning Algorithm for Path Planning in a Nondeterministic Environment"],"prefix":"10.34133","volume":"4","author":[{"given":"Junxiao","family":"Xue","sequence":"first","affiliation":[{"name":"School of Cyber Science and Engineering, Zhengzhou University, Zhengzhou, China."},{"name":"Research Center for Space Computing System, Zhejiang Lab, Hangzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-3845-5114","authenticated-orcid":true,"given":"Jinpu","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Engineering, Zhengzhou University, Zhengzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiwen","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Engineering, Zhengzhou University, Zhengzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","published-online":{"date-parts":[[2025,6,3]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1016\/j.jbusres.2020.05.019","article-title":"Robot will take your job: Innovation for an era of artificial intelligence","volume":"116","author":"Rampersad G","year":"2020","unstructured":"Rampersad G. Robot will take your job: Innovation for an era of artificial intelligence. J Bus Res. 2020;116:68\u201374.","journal-title":"J Bus Res"},{"issue":"14","key":"e_1_3_3_3_2","doi-asserted-by":"crossref","first-page":"2162","DOI":"10.3390\/electronics11142162","article-title":"A review on autonomous vehicles: Progress, methods and challenges","volume":"11","author":"Parekh D","year":"2022","unstructured":"Parekh D, Poddar N, Rajpurkar A, Chahal M, Kumar N, Joshi GP, Cho W. A review on autonomous vehicles: Progress, methods and challenges. Electronics. 2022;11(14):2162.","journal-title":"Electronics"},{"issue":"4","key":"e_1_3_3_4_2","doi-asserted-by":"crossref","first-page":"582","DOI":"10.1016\/j.dt.2019.04.011","article-title":"A review: On path planning strategies for navigation of mobile robot","volume":"15","author":"Patle BK","year":"2019","unstructured":"Patle BK, Babu LG, Pandey A, Parhi D, Jagadeesh A. A review: On path planning strategies for navigation of mobile robot. Def Technol. 2019;15(4):582\u2013606.","journal-title":"Def Technol"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.301"},{"key":"e_1_3_3_6_2","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1613\/jair.1.12440","article-title":"Reward machines: Exploiting reward function structure in reinforcement learning","volume":"73","author":"Icarte RT","year":"2022","unstructured":"Icarte RT, Klassen TQ, Valenzano R, McIlraith SA. Reward machines: Exploiting reward function structure in reinforcement learning. J Artif Intell Res. 2022;73:173\u2013208.","journal-title":"J Artif Intell Res"},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","unstructured":"Jesus JC Bottega JA Cuadros MA Gamarra DF. Deep deterministic policy gradient for navigation of mobile robots in simulated environments. In: 2019 19th International Conference on Advanced Robotics (ICAR). IEEE; 2019. p. 362\u2013367.","DOI":"10.1109\/ICAR46387.2019.8981638"},{"key":"e_1_3_3_8_2","doi-asserted-by":"crossref","unstructured":"Zhao P Zheng J Zhou Q Lyu C Lyu L. A dueling-DDPG architecture for mobile robots path planning based on laser range findings. In: PRICAI 2021: Trends in Artificial Intelligence: 18th Pacific Rim International Conference on Artificial Intelligence PRICAI 2021 Hanoi Vietnam November 8\u201312 2021 Proceedings Part I 18. Springer; 2021. p. 154\u2013168.","DOI":"10.1007\/978-3-030-89188-6_12"},{"key":"e_1_3_3_9_2","doi-asserted-by":"crossref","first-page":"1376215","DOI":"10.3389\/fnbot.2024.1376215","article-title":"Curiosity model policy optimization for robotic manipulator tracking control with input saturation in uncertain environment","volume":"18","author":"Wang T","year":"2024","unstructured":"Wang T, Wang F, Xie Z, Qin F. Curiosity model policy optimization for robotic manipulator tracking control with input saturation in uncertain environment. Front Neurorobot. 2024;18:1376215.","journal-title":"Front Neurorobot"},{"key":"e_1_3_3_10_2","doi-asserted-by":"crossref","first-page":"569","DOI":"10.1007\/s10514-022-10039-8","article-title":"Motion planning and control for mobile robot navigation using machine learning: A survey","volume":"46","author":"Xiao X","year":"2022","unstructured":"Xiao X, Liu B, Warnell G, Stone P. Motion planning and control for mobile robot navigation using machine learning: A survey. Auton Robot. 2022;46:569\u2013597.","journal-title":"Auton Robot"},{"key":"e_1_3_3_11_2","doi-asserted-by":"crossref","unstructured":"Xin J Zhao H Liu D Li M. Application of deep reinforcement learning in mobile robot path planning. In: 2017 Chinese Automation Congress (CAC). IEEE; 2017. p. 7112\u20137116.","DOI":"10.1109\/CAC.2017.8244061"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"e_1_3_3_14_2","unstructured":"Konda V Tsitsiklis J. Actor-critic algorithms. In: Advances in neural information processing systems. MIT Press; 1999. Vol. 12."},{"key":"e_1_3_3_15_2","doi-asserted-by":"crossref","unstructured":"Gu S Holly E Lillicrap T Levine S. Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In: 2017 IEEE international conference on robotics and automation (ICRA). IEEE; 2017. p. 3389\u20133396.","DOI":"10.1109\/ICRA.2017.7989385"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Tai L Paolo G Liu M. Virtual-to-real deep reinforcement learning: Continuous control of mobile robots for mapless navigation. In: 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE; 2017. p. 31\u201336.","DOI":"10.1109\/IROS.2017.8202134"},{"key":"e_1_3_3_17_2","doi-asserted-by":"crossref","first-page":"2300444","DOI":"10.1002\/aisy.202300444","article-title":"Bidirectional obstacle avoidance enhancement-deep deterministic policy gradient: A novel algorithm for mobile-robot path planning in unknown dynamic environments","volume":"6","author":"Xue J","year":"2024","unstructured":"Xue J, Zhang S, Lu Y, Yan X, Zheng Y. Bidirectional obstacle avoidance enhancement-deep deterministic policy gradient: A novel algorithm for mobile-robot path planning in unknown dynamic environments. Adv Intell Syst. 2024;6:2300444.","journal-title":"Adv Intell Syst"},{"key":"e_1_3_3_18_2","doi-asserted-by":"crossref","unstructured":"Li P Wang Y Gao Z. Path planning of mobile robot based on improved TD3 algorithm. In: 2022 IEEE International Conference on Mechatronics and Automation (ICMA). IEEE; 2022. p. 715\u2013720.","DOI":"10.1109\/ICMA54519.2022.9856399"},{"key":"e_1_3_3_19_2","unstructured":"Schaul T Quan J Antonoglou I Silver D. Prioritized experience replay. arXiv. 2015. https:\/\/doi.org\/10.48550\/arXiv.1511.05952"},{"key":"e_1_3_3_20_2","unstructured":"Fujimoto S Hoof H Meger D. Addressing function approximation error in actor-critic methods. In: International conference on machine learning. PMLR; 2018. p. 1587\u20131596."},{"issue":"9","key":"e_1_3_3_21_2","doi-asserted-by":"crossref","first-page":"3579","DOI":"10.3390\/s22093579","article-title":"Efficient path planning for mobile robot based on deep deterministic policy gradient","volume":"22","author":"Gong H","year":"2022","unstructured":"Gong H, Wang P, Ni C, Cheng N. Efficient path planning for mobile robot based on deep deterministic policy gradient. Sensors. 2022;22(9):3579.","journal-title":"Sensors"},{"key":"e_1_3_3_22_2","doi-asserted-by":"crossref","unstructured":"Zhou Q Lyu L Liu H. Deep reinforcement learning with long-time memory capability for robot mapless navigation. In: 2022 IEEE 25th International Conference on Computer Supported Cooperative Work in Design (CSCWD). IEEE; 2022. p. 1215\u20131220.","DOI":"10.1109\/CSCWD54268.2022.9776137"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco_a_01199"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","first-page":"139171","DOI":"10.1109\/ACCESS.2023.3340719","article-title":"ps-CALR: Periodic-shift cosine annealing learning rate for deep neural networks","volume":"11","author":"Johnson OV","year":"2023","unstructured":"Johnson OV, Xinying C, Khaw KW, Lee MH. ps-CALR: Periodic-shift cosine annealing learning rate for deep neural networks. IEEE Access. 2023;11:139171\u2013139186.","journal-title":"IEEE Access"},{"key":"e_1_3_3_25_2","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1016\/bs.pbr.2016.05.005","article-title":"Intrinsic motivation, curiosity, and learning: Theory and applications in educational technologies","volume":"229","author":"Oudeyer PY","year":"2016","unstructured":"Oudeyer PY, Gottlieb J, Lopes M. Intrinsic motivation, curiosity, and learning: Theory and applications in educational technologies. Prog Brain Res. 2016;229:257\u2013284.","journal-title":"Prog Brain Res"},{"key":"e_1_3_3_26_2","doi-asserted-by":"crossref","unstructured":"Pathak D Agrawal P Efros AA Darrell T. Curiosity-driven exploration by self-supervised prediction. In: International conference on machine learning. PMLR; 2017. p. 2778\u20132787.","DOI":"10.1109\/CVPRW.2017.70"},{"key":"e_1_3_3_27_2","unstructured":"Zhelo O Zhang J Tai L Liu M Burgard W. Curiosity-driven exploration for mapless navigation with deep reinforcement learning. arXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1804.00456"},{"key":"e_1_3_3_28_2","doi-asserted-by":"crossref","unstructured":"Silvia PJ. Curiosity and motivation. In: Ryan RM editor. The Oxford handbook of human motivation. New York (NY): Oxford University Press; 2012. p. 157\u2013166.","DOI":"10.1093\/oxfordhb\/9780195399820.013.0010"},{"key":"e_1_3_3_29_2","doi-asserted-by":"crossref","unstructured":"Kashdan TB Silvia PJ. Curiosity and interest: The benefits of thriving on novelty and challenge. In: Lopez SJ Snyder CR editors. The Oxford handbook of positive psychology. 2nd ed. New York (NY): Oxford University Press; 2009. p. 367\u2013374.","DOI":"10.1093\/oxfordhb\/9780195187243.013.0034"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-019-12552-4"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1177\/0959354394044004"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijpsycho.2016.11.001"},{"issue":"3","key":"e_1_3_3_33_2","doi-asserted-by":"crossref","first-page":"2244","DOI":"10.1109\/TII.2024.3495775","article-title":"Goal-conditioned reinforcement learning with adaptive intrinsic curiosity and universal value network fitting for robotic manipulation","volume":"21","author":"Sun Z","year":"2024","unstructured":"Sun Z, Yuan X, Xu Q, Pang B, Song Y, Song R, Li Y. Goal-conditioned reinforcement learning with adaptive intrinsic curiosity and universal value network fitting for robotic manipulation. IEEE Trans Industr Inform. 2024;21(3):2244\u20132253.","journal-title":"IEEE Trans Industr Inform"},{"key":"e_1_3_3_34_2","doi-asserted-by":"crossref","unstructured":"Cazenave T Sentuc J Videau M. Cosine annealing mixnet and swish activation for computer Go. In: Advances in Computer Games. Springer; 2021. p. 53\u201360.","DOI":"10.1007\/978-3-031-11488-5_5"},{"key":"e_1_3_3_35_2","doi-asserted-by":"crossref","unstructured":"Puterman ML. Markov decision processes. In: Heyman DP Sobel MJ editors. Handbooks in operations research and management science. Vol. 2. Amsterdam (Netherlands): North-Holland; 1990. p. 331\u2013434.","DOI":"10.1016\/S0927-0507(05)80172-0"},{"key":"e_1_3_3_36_2","doi-asserted-by":"crossref","unstructured":"Koenig N Howard A. Design and use paradigms for Gazebo an open-source multi-robot simulator. In: 2004 IEEE\/RSJ international conference on intelligent robots and systems (IROS) (IEEE cat. no. 04CH37566). IEEE; 2004. Vol. 3 p. 2149\u20132154.","DOI":"10.1109\/IROS.2004.1389727"},{"key":"e_1_3_3_37_2","unstructured":"Haugseter S. Autonomous driving and machine learning with TurtleBot3 Waffle Pi mobile rover [thesis]. University of South-Eastern Norway; 2023."},{"issue":"1","key":"e_1_3_3_38_2","doi-asserted-by":"crossref","first-page":"305","DOI":"10.3390\/s22010305","article-title":"Sensor data fusion for a mobile robot using a neural network algorithm","volume":"22","author":"Barreto-Cubero AJ","year":"2021","unstructured":"Barreto-Cubero AJ, G\u00f3mez-Espinosa A, Escobedo Cabello JA, Cuan-Urquizo E, Cruz-Ram\u00edrez SR. Sensor data fusion for a mobile robot using a neural network algorithm. Sensors. 2021;22(1):305.","journal-title":"Sensors"},{"key":"e_1_3_3_39_2","unstructured":"Haarnoja T Zhou A Abbeel P Levine S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In: International conference on machine learning. PMLR; 2018. p. 1861\u20131870."},{"key":"e_1_3_3_40_2","doi-asserted-by":"crossref","unstructured":"Imambi S Prakash KB Kanagachidambaresan GR. PyTorch. In: Prakash KB Kanagachidambaresan GR editors. Programming with TensorFlow: Solution for edge computing applications. Cham (Switzerland): Springer International Publishing; 2021. p. 87\u2013104.","DOI":"10.1007\/978-3-030-57077-4_10"},{"key":"e_1_3_3_41_2","first-page":"21","article-title":"Power comparisons of Shapiro-Wilk, Kolmogorov-Smirnov, Lilliefors and Anderson-Darling tests","volume":"2","author":"Razali NM","year":"2011","unstructured":"Razali NM, Wah YB. Power comparisons of Shapiro-Wilk, Kolmogorov-Smirnov, Lilliefors and Anderson-Darling tests. J Stat Model Anal. 2011;2:21\u201333.","journal-title":"J Stat Model Anal"},{"key":"e_1_3_3_42_2","doi-asserted-by":"crossref","DOI":"10.1002\/9780470479216.corpsy0491","article-title":"Kruskal-Wallis test","author":"McKight PE","year":"2010","unstructured":"McKight PE, Najab J. Kruskal-Wallis test. Corsini Encyclop Psychol. 2010; https:\/\/doi.org\/10.1002\/9780470479216.corpsy0491.","journal-title":"Corsini Encyclop Psychol"},{"issue":"2","key":"e_1_3_3_43_2","doi-asserted-by":"crossref","first-page":"99","DOI":"10.2307\/3001913","article-title":"Comparing individual means in the analysis of variance","volume":"5","author":"Tukey JW","year":"1949","unstructured":"Tukey JW. Comparing individual means in the analysis of variance. Biometrics. 1949;5(2):99\u2013114.","journal-title":"Biometrics"}],"container-title":["Intelligent Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/spj.science.org\/doi\/pdf\/10.34133\/icomputing.0140","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,3]],"date-time":"2025-06-03T15:15:57Z","timestamp":1748963757000},"score":1,"resource":{"primary":{"URL":"https:\/\/spj.science.org\/doi\/10.34133\/icomputing.0140"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1]]},"references-count":42,"alternative-id":["10.34133\/icomputing.0140"],"URL":"https:\/\/doi.org\/10.34133\/icomputing.0140","relation":{},"ISSN":["2771-5892"],"issn-type":[{"value":"2771-5892","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1]]},"assertion":[{"value":"2024-09-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-07","order":1,"name":"revised","label":"Revised","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-02","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"0140"}}