{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T06:26:56Z","timestamp":1772000816713,"version":"3.50.1"},"reference-count":30,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T00:00:00Z","timestamp":1771977600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Robot. AI"],"abstract":"<jats:p>The bipedal wheel-legged robot combines the high energy efficiency of wheeled movement with the terrain adaptability of legged locomotion. However, achieving a smooth transition between these two heterogeneous motion modes within a unified control framework remains challenging. This study proposes a reinforcement learning control framework that integrates the Mixture of Experts (MoE) architecture. This approach employs a \u201cdivide and conquer\u201d strategy by introducing a dynamic gating network and a Top-K sparse activation mechanism, which automatically allocates different motion modes to specific expert subnetworks, effectively decoupling conflicting gradients. Simulation results demonstrate that, compared to the single-network PPO method, the MoE-enhanced algorithm exhibits significant improvements in training stability and rewards. The learned policy successfully achieved smooth rolling on flat surfaces and transitioned to dynamic leg-lifting gaits when confronted with obstacles. In various test terrains, it showed a markedly higher success rate compared to the single-network PPO method.<\/jats:p>","DOI":"10.3389\/frobt.2026.1788395","type":"journal-article","created":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T05:33:20Z","timestamp":1771997600000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Adaptive multi-mode locomotion for bipedal wheel-legged robots via sparse mixture-of-experts deep reinforcement learning"],"prefix":"10.3389","volume":"13","author":[{"given":"Pan","family":"He","sequence":"first","affiliation":[{"name":"Institute of Advanced Structure Technology, Beijing Institute of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zeang","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Advanced Structure Technology, Beijing Institute of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shengyu","family":"Duan","sequence":"additional","affiliation":[{"name":"Institute of Advanced Structure Technology, Beijing Institute of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Panding","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Advanced Structure Technology, Beijing Institute of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongshuai","family":"Lei","sequence":"additional","affiliation":[{"name":"Institute of Advanced Structure Technology, Beijing Institute of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2026,2,25]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"6795","DOI":"10.1109\/TPAMI.2021.3103132","article-title":"Continuous action reinforcement learning from a mixture of interpretable experts","volume":"44","author":"Akrour","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"B2","doi-asserted-by":"publisher","first-page":"2354","DOI":"10.1109\/TMECH.2020.2973752","article-title":"Design and control of a compliant wheel-on-leg rover which conforms to uneven terrain","volume":"25","author":"Bouton","year":"2020","journal-title":"IEEE\/ASME Trans. Mechatronics"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2403.06966","article-title":"Acquiring diverse skills using curriculum reinforcement learning with mixture of experts","author":"Celik","year":"2024"},{"key":"B4","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2503.07049","article-title":"VMTS: vision-assisted teacher-student reinforcement learning for multi-terrain locomotion in bipedal robots","author":"Chen","year":"2025"},{"key":"B5","doi-asserted-by":"publisher","first-page":"7667","DOI":"10.1109\/LRA.2021.3100269","article-title":"Learning-based balance control of wheel-legged robots","volume":"6","author":"Cui","year":"2021","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B6","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2511.12361","article-title":"SAC-MoE: reinforcement learning with mixture-of-experts for control of hybrid dynamical systems with uncertainty","author":"D\u2019Souza","year":"2025"},{"key":"B7","doi-asserted-by":"publisher","first-page":"100283","DOI":"10.1016\/j.finmec.2024.100283","article-title":"Mobile rolling robots designed to overcome obstacles: a review","volume":"16","author":"Garc\u00eda","year":"2024","journal-title":"Forces Mech."},{"key":"B8","doi-asserted-by":"publisher","first-page":"1066714","DOI":"10.3389\/fnbot.2022.1066714","article-title":"Design and dynamic analysis of jumping wheel-legged robot in complex terrain environment","volume":"16","author":"Guo","year":"2022","journal-title":"Front. Neurorobot."},{"key":"B9","doi-asserted-by":"publisher","DOI":"10.20517\/ss.2024.50","article-title":"A review: exploring the designs of bio-bots","author":"He","year":"2025","journal-title":"Soft Sci."},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2410.14972","article-title":"MENTOR: mixture-of-experts network with task-oriented perturbation for visual reinforcement learning","author":"Huang","year":"2025"},{"key":"B11","doi-asserted-by":"publisher","first-page":"eaau5872","DOI":"10.1126\/scirobotics.aau5872","article-title":"Learning agile and dynamic motor skills for legged robots","volume":"4","author":"Hwangbo","year":"2019","journal-title":"Sci. Robot."},{"key":"B30","doi-asserted-by":"publisher","first-page":"3521","DOI":"10.1073\/pnas.1611835114","article-title":"Overcoming catastrophic forgetting in neural networks","volume":"114","author":"Kirkpatrick","year":"2017","journal-title":"Proc. Natl. Acad. Sci."},{"key":"B12","doi-asserted-by":"crossref","first-page":"7515","DOI":"10.1109\/ICRA.2019.8793792","article-title":"Ascento: a two-wheeled jumping robot","volume-title":"2019 international conference on robotics and automation (ICRA)","author":"Klemm","year":"2019"},{"key":"B13","doi-asserted-by":"publisher","first-page":"3745","DOI":"10.1109\/LRA.2020.2979625","article-title":"LQR-assisted whole-body control of a wheeled bipedal robot with kinematic loops","volume":"5","author":"Klemm","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2107.04034","article-title":"RMA: rapid motor adaptation for legged robots","author":"Kumar","year":"2021"},{"key":"B15","doi-asserted-by":"publisher","first-page":"eabc5986","DOI":"10.1126\/scirobotics.abc5986","article-title":"Learning quadrupedal locomotion over challenging terrain","volume":"5","author":"Lee","year":"2020","journal-title":"Sci. Robot."},{"key":"B16","doi-asserted-by":"publisher","first-page":"eadi9641","DOI":"10.1126\/scirobotics.adi9641","article-title":"Learning robust autonomous navigation and locomotion for wheeled-legged robots","volume":"9","author":"Lee","year":"2024","journal-title":"Sci. Robot."},{"key":"B17","doi-asserted-by":"crossref","first-page":"669","DOI":"10.1145\/3677052.3698691","article-title":"Mixtures of experts for scaling up neural networks in order execution","volume-title":"Proceedings of the 5th ACM international conference on AI in finance","author":"Li","year":"2024"},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2301.10602","article-title":"DreamWaQ: learning robust quadrupedal locomotion with implicit terrain imagination via deep reinforcement learning","author":"Nahrendra","year":"2023"},{"key":"B19","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.08609","article-title":"Mixtures of experts unlock parameter scaling for deep RL","author":"Obando-Ceron","year":"2024"},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1904.11455","article-title":"Ray interference: a source of plateaus in deep reinforcement learning","author":"Schaul","year":"2019"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1707.06347","article-title":"Proximal policy optimization algorithms","author":"Schulman","year":"2017"},{"key":"B22","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1701.06538","article-title":"Outrageously large neural networks: the sparsely-gated mixture-of-experts layer","author":"Shazeer","year":"2017"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.1109\/Humanoids57100.2023.10375214","article-title":"Analytical second-order derivatives of rigid-body contact dynamics: application to multi-shooting DDP","author":"Singh","year":""},{"key":"B24","doi-asserted-by":"publisher","first-page":"20982","DOI":"10.1609\/aaai.v39i20.35394","article-title":"SMOSE: sparse mixture of shallow experts for interpretable reinforcement learning in continuous control tasks","volume":"39","author":"Vincze","year":"2025","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"B25","first-page":"6782","article-title":"Balance control of a novel wheel-legged robot: design and experiments","volume-title":"2021 IEEE international conference on robotics and automation (ICRA), (Xi\u2019an, China: IEEE)","author":"Wang","year":"2021"},{"key":"B26","doi-asserted-by":"publisher","first-page":"100256","DOI":"10.1016\/j.birob.2025.100256","article-title":"Wheeled-legged robots for multi-terrain locomotion in plateau environments","volume":"5","author":"Wang","year":"2025","journal-title":"Biomim. Intell. Rob."},{"key":"B27","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2507.07818","article-title":"MoSE: skill-by-skill mixture-of-experts learning for embodied autonomous machines","author":"Xu","year":"2025"},{"key":"B28","article-title":"Gradient surgery for multi-task learning","author":"Yu","year":"2020"},{"key":"B29","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TMECH.2024.3522904","article-title":"Tensegrity-based legged robot generates passive walking, skipping, and crawling gaits in accordance with environment","volume":"30","author":"Zheng","year":"2025","journal-title":"IEEE\/ASME Trans. Mechatronics"}],"container-title":["Frontiers in Robotics and AI"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frobt.2026.1788395\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T05:33:23Z","timestamp":1771997603000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frobt.2026.1788395\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,25]]},"references-count":30,"alternative-id":["10.3389\/frobt.2026.1788395"],"URL":"https:\/\/doi.org\/10.3389\/frobt.2026.1788395","relation":{},"ISSN":["2296-9144"],"issn-type":[{"value":"2296-9144","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,25]]},"article-number":"1788395"}}