{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T04:08:58Z","timestamp":1778472538357,"version":"3.51.4"},"reference-count":37,"publisher":"Cambridge University Press (CUP)","issue":"2","license":[{"start":{"date-parts":[[2026,2,18]],"date-time":"2026-02-18T00:00:00Z","timestamp":1771372800000},"content-version":"unspecified","delay-in-days":17,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Robotica"],"published-print":{"date-parts":[[2026,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>With the rapid development of artificial intelligence technology, robotics, as its core branch, has attracted extensive attention from researchers. This paper designs and develops a robotic arm learning system based on multi-source sensor information fusion, which investigates the autonomous learning capability of robotic arms by closely integrating deep reinforcement learning (DRL) as the core framework for skill acquisition. By incorporating imitation learning as a source of expert prior and leveraging DRL\u2019s intrinsic ability for policy optimization through environmental exploration, the proposed system achieves both rapid learning and robust generalization. Specifically, we introduce the gradient penalty mechanism from Wasserstein generative adversarial networks (WGANs), a technique that improves the stability of adversarial training by penalizing gradients that deviate from a specified norm. This mechanism is incorporated into the soft actor-critic (SAC) algorithm, a widely used off-policy DRL method known for its sample efficiency and robust performance in continuous control tasks. The resulting SAC-GP (SAC-gradient penalty) algorithm benefits from both SAC\u2019s stable policy learning and WGAN\u2019s improved training regularization, leading to superior convergence speed and system stability. Furthermore, this paper proposes a hybrid learning framework by combining generative adversarial imitation learning (GAIL) with SAC-GP, enabling the agent to benefit from both demonstration-based policy initialization and continuous self-improvement via reinforcement learning. Finally, a door-opening experiment is designed to verify the learning and execution capabilities of the system in both virtual and real environments. Experimental results demonstrate that the proposed learning system possesses excellent learning and motion execution abilities in practical applications. This achievement not only provides new insights for research in robot learning but also lays a solid foundation for the future development of robotic technology.<\/jats:p>","DOI":"10.1017\/s0263574726103191","type":"journal-article","created":{"date-parts":[[2026,2,18]],"date-time":"2026-02-18T06:39:41Z","timestamp":1771396781000},"page":"540-564","source":"Crossref","is-referenced-by-count":0,"title":["Robot autonomous learning system based on multi-source sensor information fusion"],"prefix":"10.1017","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6908-1580","authenticated-orcid":false,"given":"Ze","family":"Cui","sequence":"first","affiliation":[{"id":[{"id":"https:\/\/ror.org\/006teas31","id-type":"ROR","asserted-by":"publisher"}],"name":"Shanghai University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chenzhao","family":"Sun","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/006teas31","id-type":"ROR","asserted-by":"publisher"}],"name":"Shanghai University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zenghao","family":"Chen","sequence":"additional","affiliation":[{"name":"Shanghai Aerospace Control Technology Institute"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jing\u2019ao","family":"Huang","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/006teas31","id-type":"ROR","asserted-by":"publisher"}],"name":"Shanghai University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Zhang","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/006teas31","id-type":"ROR","asserted-by":"publisher"}],"name":"Shanghai University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2026,2,18]]},"reference":[{"key":"S0263574726103191_ref29","doi-asserted-by":"crossref","unstructured":"[29] Azzam, R. , Chehadeh, M. , Abdulhay, O. , Boiko, I. and Zweiri, Y. , \u201cLearning to navigate through reinforcement across the Sim2Real gap,\u201d TechRxiv, (2022).10.36227\/techrxiv.20138960.v1","DOI":"10.36227\/techrxiv.20138960"},{"key":"S0263574726103191_ref35","doi-asserted-by":"publisher","DOI":"10.1016\/j.rcim.2023.102570"},{"key":"S0263574726103191_ref30","unstructured":"[30] Abeyruwan, S. , Graesser, L. , D\u2019Ambrosio, D. B. , Singh, A. , Shankar, A. , Bewley, A. , Jain, D. , Choromanski, K. and Sanketi, P. R. , \u201cI-Sim2Real: Reinforcement learning of robotic policies in tight human\u2013robot interaction loops, arXiv preprint arXiv: 2207.06572, 2022."},{"key":"S0263574726103191_ref9","unstructured":"[9] Xu, C. , Zhao, Y.-B. , Lu, Z. and Zhang, Y. , \u201cReinforcement-learning-based algorithms for optimization problems and applications to inverse problems,\u201d (2025)."},{"key":"S0263574726103191_ref13","unstructured":"[13] Haarnoja, T. , Zhou, A. , Hartikainen, K. , Tucker, G. , Ha, S. , Tan, J. , Kumar, V. , Zhu, H. , Gupta, A. , Abbeel, P. and Levine, S. , \u201cSoft actor-critic algorithms and applications, arXiv preprint arXiv: 1812.05905, 2018."},{"key":"S0263574726103191_ref28","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3292075"},{"key":"S0263574726103191_ref17","doi-asserted-by":"publisher","DOI":"10.1109\/TAI.2023.3328848"},{"key":"S0263574726103191_ref22","first-page":"19","volume-title":"Proceedings of the 25th International Conference on Neural Information Processing Systems (NIPS), NIPS\u201911","author":"Levine","year":"2011"},{"key":"S0263574726103191_ref1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jmsy.2024.03.004"},{"key":"S0263574726103191_ref10","doi-asserted-by":"publisher","DOI":"10.3390\/s24020561"},{"key":"S0263574726103191_ref11","doi-asserted-by":"crossref","unstructured":"[11] Lyu, L. , Shen, Y. and Zhang, S. , \u201cThe Advance of Reinforcement Learning and Deep Reinforcement Learning,\u201d In: 2022 IEEE International Conference on Electrical Engineering, Big Data and Algorithms (EEBDA) (2022) pp. 644\u2013648.","DOI":"10.1109\/EEBDA53927.2022.9744760"},{"key":"S0263574726103191_ref6","doi-asserted-by":"publisher","DOI":"10.1145\/3443686"},{"key":"S0263574726103191_ref23","unstructured":"[23] Gulrajani, I. , Ahmed, F. , Arjovsky, M. , Dumoulin, V. and Courville, A. , \u201cImproved training of wasserstein gans. arXiv preprint arXiv: 1704.00028, 2017."},{"key":"S0263574726103191_ref18","doi-asserted-by":"crossref","unstructured":"[18] Huang, T. , Chen, K. , Li, B. , Liu, Y.-H. and Dou, Q. , \u201cDemonstration-guided reinforcement learning with efficient exploration for task automation of surgical robot. arXiv preprint arXiv: 2302.09772, 2023.","DOI":"10.1109\/ICRA48891.2023.10160327"},{"key":"S0263574726103191_ref4","first-page":"20","article-title":"AI-based self-learning robotic arm using microcontroller","volume":"10","author":"Kumar","year":"2024","journal-title":"J. Control Instrum. Eng."},{"key":"S0263574726103191_ref32","first-page":"1066","article-title":"Sim2Real Transfer for Deep Reinforcement Learning with Stochastic State Transition Delays","volume":"155","author":"Sandha","year":"2021","journal-title":"Proceedings of the 2020 Conference on Robot Learning"},{"key":"S0263574726103191_ref34","doi-asserted-by":"publisher","DOI":"10.1109\/TMECH.2023.3330427"},{"key":"S0263574726103191_ref2","doi-asserted-by":"publisher","DOI":"10.1007\/s00464-023-10554-4"},{"key":"S0263574726103191_ref12","unstructured":"[12] Haarnoja, T. , Zhou, A. , Abbeel, P. and Levine, S. , \u201cSoft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, arXiv preprint arXiv: 1801.01290, 2018."},{"key":"S0263574726103191_ref7","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.106945"},{"key":"S0263574726103191_ref37","doi-asserted-by":"crossref","unstructured":"[37] Gupta, A. , Yu, J. , Zhao, T. Z. , Kumar, V. , Rovinsky, A. , Xu, K. , Devlin, T. and Levine, S. , \u201cReset-Free Reinforcement Learning Via Multi-task Learning: Learning Dexterous Manipulation Behaviors Without Human Intervention,\u201d In: 2021 IEEE International Conference on Robotics and Automation (ICRA) (2021) pp. 6664\u20136671.","DOI":"10.1109\/ICRA48506.2021.9561384"},{"key":"S0263574726103191_ref3","doi-asserted-by":"publisher","DOI":"10.1016\/j.compag.2024.108791"},{"key":"S0263574726103191_ref26","doi-asserted-by":"crossref","unstructured":"[26] Todorov, E. , Erez, T. and Tassa, Y. , \u201cMujoco: A Physics Engine for Model-based Control,\u201d In: 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (2012) pp. 5026\u20135033.","DOI":"10.1109\/IROS.2012.6386109"},{"key":"S0263574726103191_ref19","doi-asserted-by":"crossref","unstructured":"[19] Zhu, Y. , Wang, Z. , Merel, J. , Rusu, A. , Erez, T. , Cabi, S. , Tunyasuvunakool, S. , Kram\u00e1r, J. , Hadsell, R. , de Freitas, N. and Heess, N. , \u201cReinforcement and imitation learning for diverse visuomotor skills. arXiv preprint arXiv: 1802.09564, 2018.","DOI":"10.15607\/RSS.2018.XIV.009"},{"key":"S0263574726103191_ref25","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-control-053018-023825"},{"key":"S0263574726103191_ref8","doi-asserted-by":"publisher","DOI":"10.26599\/TST.2021.9010012"},{"key":"S0263574726103191_ref27","unstructured":"[27] Zhu, Y. , Wong, J. , Mandlekar, A. , Mart\u00edn-Mart\u00edn, R. , Joshi, A. , Nasiriany, S. and Zhu, Y. , \u201cRobosuite: A modular simulation framework and benchmark for robot learning. arXiv preprint arXiv: 2009.12293, 2020."},{"key":"S0263574726103191_ref16","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2025.129799"},{"key":"S0263574726103191_ref24","unstructured":"[24] Sutton, R. S. and Barto, A. G. , Reinforcement Learning: An Introduction (MIT Press, Cambridge, MA, 1998).10.1109\/TNN.1998.712192"},{"key":"S0263574726103191_ref31","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3303714"},{"key":"S0263574726103191_ref15","doi-asserted-by":"publisher","DOI":"10.1109\/TASE.2025.3594680"},{"key":"S0263574726103191_ref20","unstructured":"[20] Wu, P. , Escontrela, A. , Hafner, D. , Goldberg, K. and Abbeel, P. , \u201cDaydreamer: World models for physical robot learning. arXiv preprint arXiv: 2206.14176, 2022."},{"key":"S0263574726103191_ref36","doi-asserted-by":"publisher","DOI":"10.1109\/TASE.2023.3310945"},{"key":"S0263574726103191_ref5","unstructured":"[5] Kroemer, O. , Niekum, S. and Konidaris, G. , \u201cA review of robot learning for manipulation: Challenges, representations, and algorithms. arXiv preprint arXiv: 1907.03146, 2019."},{"key":"S0263574726103191_ref33","doi-asserted-by":"publisher","DOI":"10.1109\/TASE.2023.3288037"},{"key":"S0263574726103191_ref21","doi-asserted-by":"publisher","DOI":"10.1007\/s41315-022-00236-0"},{"key":"S0263574726103191_ref14","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2023.106613"}],"container-title":["Robotica"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S0263574726103191","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T03:32:25Z","timestamp":1778470345000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S0263574726103191\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2]]},"references-count":37,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,2]]}},"alternative-id":["S0263574726103191"],"URL":"https:\/\/doi.org\/10.1017\/s0263574726103191","relation":{},"ISSN":["0263-5747","1469-8668"],"issn-type":[{"value":"0263-5747","type":"print"},{"value":"1469-8668","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2]]}}}