{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T06:52:01Z","timestamp":1777704721494,"version":"3.51.4"},"reference-count":39,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2018,8,2]],"date-time":"2018-08-02T00:00:00Z","timestamp":1533168000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"published-print":{"date-parts":[[2019,2,16]]},"abstract":"<jats:p>Passive dynamic walking (PDW) has attracted much research attention due to its humanoid and energy efficient gaits. However, walking control of the passivity-based biped robot inspired by PDW still remains a challenge, for PDW is sensitive to disturbances. An walking controller is essential for practical passivity-based biped robots in real environments. This paper presents a deep reinforcement learning (DRL) controller based on deep Q network for planar passivity-based biped robot, to learn policies directly from inputs for bipedal walking task. First, the intelligent controller using deep Q network is trained, with PDW as reference trajectory. The learning experience from PDW could be helpful to implement a natural looking and energy-efficient gait. Then the trained deep Q network is utilized as the walking controller. Simulation results show that the DRL controller based on deep Q network makes the planar biped robot walk against original value disturbance, on different slope, level ground and varying slopes. The controller this paper presented could be used to improve the versatility of the passivity-based biped robot.<\/jats:p>","DOI":"10.3233\/jifs-172180","type":"journal-article","created":{"date-parts":[[2018,8,5]],"date-time":"2018-08-05T06:39:53Z","timestamp":1533451193000},"page":"731-745","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":16,"title":["Intelligent controller for passivity-based biped robot using deep Q network"],"prefix":"10.1177","volume":"36","author":[{"given":"Yao","family":"Wu","sequence":"first","affiliation":[{"name":"School of Power and Mechanical Engineering, Wuhan University, Wuhan, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daojin","family":"Yao","sequence":"additional","affiliation":[{"name":"School of Power and Mechanical Engineering, Wuhan University, Wuhan, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaohui","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of Power and Mechanical Engineering, Wuhan University, Wuhan, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhao","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Power and Mechanical Engineering, Wuhan University, Wuhan, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2018,8,2]]},"reference":[{"key":"e_1_3_3_2_2","first-page":"2478","article-title":"The intelligent ASIMO: System overview and integration","author":"Sakagami Y.","year":"2002","unstructured":"SakagamiY., WatanabeR., AoyamaC., MatsunagaS. and FujimuraK., The intelligent ASIMO: System overview and integration, In Intelligent Robots and Systems, 2002, IEEE\/RSJ International Conference on, 2002, pp. 2478\u20132483.","journal-title":"In Intelligent Robots and Systems, 2002, IEEE\/RSJ International Conference on"},{"key":"e_1_3_3_3_2","first-page":"4400","article-title":"Humanoid robot HRP-4 - Humanoid robotics platform with lightweight and slim body","author":"Kaneko K.","year":"2011","unstructured":"KanekoK., KanehiroF., MorisawaM. and AkachiK., Humanoid robot HRP-4 - Humanoid robotics platform with lightweight and slim body, In Intelligent Robots and Systems (IROS), 2011 IEEE\/RSJ International Conference on, 2011, pp. 4400\u20134407.","journal-title":"In Intelligent Robots and Systems (IROS), 2011 IEEE\/RSJ International Conference on"},{"key":"e_1_3_3_4_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/0025-5564(72)90061-2","article-title":"On the stability of anthropomorphic systems","volume":"15","author":"Vukobratovi\u0107 M.","year":"1972","unstructured":"Vukobratovi\u0107M. and StepanenkoJ., On the stability of anthropomorphic systems, Mathematical Biosciences 15 (1972), 1\u201337.","journal-title":"Mathematical Biosciences"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.1107799"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12206-018-0336-0"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1177\/027836499000900206"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/MRA.2007.380638"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2012.2186565"},{"key":"e_1_3_3_10_2","unstructured":"SuttonR.S. BartoA.G. Reinforcement learning: An introduction Cambridge: MIT Press 1998."},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"e_1_3_3_12_2","volume-title":"On-line Q-learning using connectionist systems","author":"Rummery G.A.","year":"1994","unstructured":"RummeryG.A., NiranjanM., On-line Q-learning using connectionist systems, University of Cambridge, Department of Engineering 1994."},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00114731"},{"key":"e_1_3_3_14_2","first-page":"4173","article-title":"Reinforcement learning control for biped robot walking on uneven surfaces","author":"Wang S.","year":"2006","unstructured":"WangS., BraaksmaJ., BabuskaR. and HobbelenD., Reinforcement learning control for biped robot walking on uneven surfaces, In Neural Networks, 2006, IJCNN\u201906, International Joint Conference on, 2006, pp. 4173\u20134178.","journal-title":"In Neural Networks, 2006, IJCNN\u201906, International Joint Conference on"},{"key":"e_1_3_3_15_2","first-page":"115","article-title":"Learning a model-free robotic continuous state-action task through contractive Q-network","author":"Davari M.","year":"2017","unstructured":"DavariM., AlipourK., HadiA. and TarvirdizadehB., Learning a model-free robotic continuous state-action task through contractive Q-network, In Artificial Intelligence and Robotics (IRANOPEN) (2017), 115\u2013120.","journal-title":"In Artificial Intelligence and Robotics (IRANOPEN)"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.3233\/IFS-162212"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.3233\/JIFS-161822"},{"key":"e_1_3_3_18_2","first-page":"2849","article-title":"Stochastic policy gradient reinforcement learning on a simple 3D biped","author":"Tedrake R.","year":"2004","unstructured":"TedrakeR., ZhangT.W. and SeungH.S., Stochastic policy gradient reinforcement learning on a simple 3D biped, Intelligent Robots and Systems, 2004 (IROS 2004) Proceedings 2004 IEEE\/RSJ International Conference on, 2004, pp. 2849\u20132854.","journal-title":"Intelligent Robots and Systems, 2004 (IROS 2004) Proceedings 2004 IEEE\/RSJ International Conference on"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ifacol.2016.07.994"},{"key":"e_1_3_3_20_2","first-page":"1","article-title":"A novel approach to locomotion learning: Actor-Critic architecture using central pattern generators and dynamic motor primitives","volume":"8","author":"Li C.","year":"2014","unstructured":"LiC., LoweR. and ZiemkeT., A novel approach to locomotion learning: Actor-Critic architecture using central pattern generators and dynamic motor primitives, Frontiers in Neurorobotics 8 (2014), 1\u201317.","journal-title":"Frontiers in Neurorobotics"},{"key":"e_1_3_3_21_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/2200000006","article-title":"Learning deep architectures for AI","volume":"2","author":"Bengio Y.","year":"2009","unstructured":"BengioY., Learning deep architectures for AI, Foundations & Trends in Machine Learning 2 (2009), 1\u2013127.","journal-title":"Foundations & Trends in Machine Learning"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2743240"},{"key":"e_1_3_3_23_2","unstructured":"MnihV. KavukcuogluK. SilverD. GravesA. AntonoglouI. WierstraD. and RiedmillerM. Playing Atari with Deep Reinforcement Learning 2013 arXiv:1312.5602."},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.3233\/JIFS-169151"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature24270"},{"key":"e_1_3_3_27_2","unstructured":"LiY. Deep reinforcement learning: An overview 2017 arXiv: 1701.07274."},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_3_29_2","first-page":"2094","article-title":"Deep Reinforcement Learning with Double Q-Learning","author":"Van Hasselt H.","year":"2016","unstructured":"Van HasseltH., GuezA. and SilverD., Deep Reinforcement Learning with Double Q-Learning, In AAAI (2016), 2094\u20132100.","journal-title":"In AAAI"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1177\/027836499801701202"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.5772\/62687"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1017\/S0263574715000077"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jtbi.2007.05.008"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1017\/S0263574708004906"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1201\/9781420053739"},{"key":"e_1_3_3_36_2","volume-title":"Finite element procedures in engineering analysis","author":"Bathe K.J.","year":"1982","unstructured":"BatheK.J., Finite element procedures in engineering analysis, Prentice-Hall, 1982."},{"key":"e_1_3_3_37_2","volume-title":"Advanced contact dynamics","author":"Eberhard P.","year":"2003","unstructured":"EberhardP., HuB., Advanced contact dynamics, Southeast University Press, 2003."},{"key":"e_1_3_3_38_2","article-title":"Carnegie-Mellon Univ Pittsburgh PA School of Computer Science","author":"Lin L.J.","year":"1993","unstructured":"LinL.J., Carnegie-Mellon Univ Pittsburgh PA School of Computer Science, Reinforcement learning for robots using neural networks, 1993.","journal-title":"Reinforcement learning for robots using neural networks"},{"key":"e_1_3_3_39_2","unstructured":"AbadiM. AgarwalA. BarhamP. BrevdoE. ChenZ. CitroC. and GhemawatS. Tensorflow: Large-scale machine learning on heterogeneous distributed systems 2016 arXiv: 1603.04467."},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.egypro.2012.02.296"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-172180","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/JIFS-172180","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-172180","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:41:51Z","timestamp":1777455711000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/JIFS-172180"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,8,2]]},"references-count":39,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2019,2,16]]}},"alternative-id":["10.3233\/JIFS-172180"],"URL":"https:\/\/doi.org\/10.3233\/jifs-172180","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,8,2]]}}}