{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T18:20:58Z","timestamp":1783966858356,"version":"3.55.0"},"reference-count":31,"publisher":"Cambridge University Press (CUP)","issue":"1","license":[{"start":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T00:00:00Z","timestamp":1768262400000},"content-version":"unspecified","delay-in-days":12,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Robotica"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Legged robots have demonstrated remarkable potential for dynamic locomotion and terrain adaptability, making them a prominent focus of research. However, achieving robust and agile bipedal running remains challenging due to the complex dynamics of legged locomotion. In this paper, we propose a reinforcement learning framework for robust bipedal running, incorporating a simple reference trajectory generator and an asymmetric actor-critic architecture. The reference generator, based on kinematics, provides diverse trajectory references while preserving key gait characteristics, facilitating efficient policy exploration. To mitigate the simulation-to-reality gap, we extract latent variables encoding environmental and motion information from dual historical observations. Our method simplifies the trajectory generation process while maintaining effective guidance for learning. Extensive simulation and physical experiments demonstrate that, compared to model-based and learning-based baselines, our approach achieves higher agility, more accurate velocity tracking, and stronger disturbance rejection while preserving gait stability. The resulting controller exhibits spring\u2013mass running dynamics that remain robust on both flat and uneven terrains.<\/jats:p>","DOI":"10.1017\/s0263574725103007","type":"journal-article","created":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T08:26:49Z","timestamp":1768292809000},"page":"150-168","source":"Crossref","is-referenced-by-count":1,"title":["Learning robust bipedal running via structured gait and trajectory guidance"],"prefix":"10.1017","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-3907-8902","authenticated-orcid":false,"given":"Yunpeng","family":"Liang","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhihui","family":"Peng","sequence":"additional","affiliation":[{"name":"AgiBot Technology Co. Ltd."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanzheng","family":"Zhao","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7173-4085","authenticated-orcid":false,"given":"Weixin","family":"Yan","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"56","published-online":{"date-parts":[[2026,1,13]]},"reference":[{"key":"S0263574725103007_ref24","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3151396"},{"key":"S0263574725103007_ref26","doi-asserted-by":"crossref","unstructured":"[26] Todorov, E. , Erez, T. and Tassa, Y. , \u201cMuJoCo: A Physics Engine for Model-based Control,\u201d In: Proceedings of IEEE\/RSJ International Conference on Intelligent Robots and Systems (2012) pp. 5026\u20135033.","DOI":"10.1109\/IROS.2012.6386109"},{"key":"S0263574725103007_ref27","unstructured":"[27] Schulman, J. , Wolski, F. , Dhariwal, P. , Radford, A. and Klimov, O. , \u201cProximal policy optimization algorithms,\u201d arXiv preprint arXiv:1707.06347 (2017)."},{"key":"S0263574725103007_ref9","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3342668"},{"key":"S0263574725103007_ref13","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2017.2783371"},{"key":"S0263574725103007_ref14","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.adi9641"},{"key":"S0263574725103007_ref11","unstructured":"[11] Yang, W. and Posa, M. , \u201cImpact-invariant control: Maximizing control authority during impacts,\u201d arXiv preprint arXiv:2303.00817 (2023)."},{"key":"S0263574725103007_ref17","unstructured":"[17] Xie, Z. , Clary, P. , Dao, J. , Morais, P. , Hurst, J. and Panne, M. , \u201cLearning Locomotion Skills for Cassie: Iterative Design and Sim-to-Real,\u201d In: Conference on Robot Learning (PMLR, 2020) pp. 317\u2013329."},{"key":"S0263574725103007_ref30","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3068908"},{"key":"S0263574725103007_ref22","doi-asserted-by":"crossref","unstructured":"[22] Wu, Q. , Zhang, C. and Liu, Y. , \u201cCustom Sine Waves are Enough for Imitation Learning of Bipedal Gaits with Different Styles,\u201d In: 2022 IEEE International Conference on Mechatronics and Automation (ICMA) (2022) pp. 499\u2013505.","DOI":"10.1109\/ICMA54519.2022.9856382"},{"key":"S0263574725103007_ref10","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2024.3524902"},{"key":"S0263574725103007_ref12","doi-asserted-by":"publisher","DOI":"10.1177\/0278364912473344"},{"key":"S0263574725103007_ref5","doi-asserted-by":"publisher","DOI":"10.1017\/S0263574724001097"},{"key":"S0263574725103007_ref3","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2021.XVII.061"},{"key":"S0263574725103007_ref31","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2025.3614078"},{"key":"S0263574725103007_ref2","doi-asserted-by":"publisher","DOI":"10.23919\/ACC.2019.8814833"},{"key":"S0263574725103007_ref19","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9560769"},{"key":"S0263574725103007_ref25","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.aau5872"},{"key":"S0263574725103007_ref6","doi-asserted-by":"publisher","DOI":"10.1177\/02783649241285161"},{"key":"S0263574725103007_ref29","unstructured":"[29] Gu, X. , Wang, Y.-J. and Chen, J. , \u201cHumanoid-gym: Reinforcement learning for humanoid robot with zero-shot sim2real transfer,\u201d arXiv preprint arXiv:2404.05695 (2024)."},{"key":"S0263574725103007_ref7","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3324580"},{"key":"S0263574725103007_ref8","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2013.6697099"},{"key":"S0263574725103007_ref28","unstructured":"[28] Makoviychuk, V. , Wawrzyniak, L. , Guo, Y. , Lu, M. , Storey, K. , Macklin, M. , Hoeller, D. , Rudin, N. , Allshire, A. , Handa, A. and State, G. , \u201cIsaac gym: High performance GPU-based physics simulation for robot learning,\u201d arXiv:2108.10470 (2021)."},{"key":"S0263574725103007_ref18","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3066833"},{"key":"S0263574725103007_ref23","doi-asserted-by":"publisher","DOI":"10.1109\/MEX.1986.4307016"},{"key":"S0263574725103007_ref4","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abc5986"},{"key":"S0263574725103007_ref1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3275384"},{"key":"S0263574725103007_ref21","doi-asserted-by":"publisher","DOI":"10.1109\/IROS60139.2025.11247685"},{"key":"S0263574725103007_ref16","doi-asserted-by":"crossref","unstructured":"[16] Xie, Z. , Berseth, G. , Clary, P. , Hurst, J. and van de Panne, M. , \u201cFeedback Control for Cassie with Deep Reinforcement Learning,\u201d In: 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), vol. 1241 (2018) pp. 1246.","DOI":"10.1109\/IROS.2018.8593722"},{"key":"S0263574725103007_ref20","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9561814"},{"key":"S0263574725103007_ref15","first-page":"1","article-title":"Deepmimic: Example-guided deep reinforcement learning of physics-based character skills","volume":"37","author":"Peng","year":"2018","journal-title":"ACM Trans. Graph. (TOG)"}],"container-title":["Robotica"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S0263574725103007","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T05:02:12Z","timestamp":1780462932000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S0263574725103007\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1]]},"references-count":31,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["S0263574725103007"],"URL":"https:\/\/doi.org\/10.1017\/s0263574725103007","relation":{},"ISSN":["0263-5747","1469-8668"],"issn-type":[{"value":"0263-5747","type":"print"},{"value":"1469-8668","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1]]}}}