{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T19:54:58Z","timestamp":1780430098584,"version":"3.54.1"},"reference-count":41,"publisher":"Cambridge University Press (CUP)","issue":"8","license":[{"start":{"date-parts":[[2025,7,18]],"date-time":"2025-07-18T00:00:00Z","timestamp":1752796800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Robotica"],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The automation of assembly operations with industrial robots is pivotal in modern manufacturing, particularly for multispecies, low-volume, and customized production. Traditional programing methods are time-consuming and lack adaptability to complex, variable environments. Reinforcement learning-based assembly tasks have shown success in simulation environments, but face challenges like the simulation-to-reality gap and safety concerns when transferred to real-world applications. This article addresses these challenges by proposing a low-cost, image-segmentation-driven deep reinforcement learning strategy tailored for insertion tasks, such as the assembly of peg-in-hole components in satellite manufacturing, which involve extensive contact interactions. Our approach integrates visual and forces feedback into a prior dueling deep Q-network for insertion skill learning, enabling precise alignment of components. To bridge the simulation-to-reality gap, we transform the raw image input space into a canonical space based on image segmentation. Specifically, we employ a segmentation model based on U-net, pretrained in simulation and fine-tuned with real-world data, significantly reducing the need for labor-intensive real image segment labels. To handle the frequent contact inherent in peg-in-hole tasks, we integrated safety protections and impedance control into the training process, providing active compliance and reducing the risk of assembly failures. Our approach was evaluated in both simulated and real robotic environments, demonstrating robust performance in handling camera position errors and varying ambient light intensities and different lighting colors. Finally, the algorithm was validated in a real satellite assembly scenario, achieving a success rate of 15 out of 20 tests.<\/jats:p>","DOI":"10.1017\/s026357472510177x","type":"journal-article","created":{"date-parts":[[2025,7,18]],"date-time":"2025-07-18T07:06:36Z","timestamp":1752822396000},"page":"2783-2802","source":"Crossref","is-referenced-by-count":1,"title":["Image segmentation-driven sim-to-real deep reinforcement learning framework for accurate peg-in-hole assembly"],"prefix":"10.1017","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0127-1334","authenticated-orcid":false,"given":"Ning","family":"Zhang","sequence":"first","affiliation":[{"name":"Beihang University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yongjia","family":"Zhao","sequence":"additional","affiliation":[{"name":"Beihang University"},{"name":"Beihang University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minghao","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuling","family":"Dai","sequence":"additional","affiliation":[{"name":"Beihang University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"56","published-online":{"date-parts":[[2025,7,18]]},"reference":[{"key":"S026357472510177X_ref40","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3061374"},{"key":"S026357472510177X_ref28","doi-asserted-by":"crossref","unstructured":"[28] Stemmer, A. , Schreiber, G. , Arbter, K. and Albu-Schaffer, A. , \u201cRobust Assembly of Complex Shaped Planar Parts Using Vision And Force,\u201d 2006 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems, IEEE (2006) pp. 493\u2013500.","DOI":"10.1109\/MFI.2006.265671"},{"key":"S026357472510177X_ref2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.3010739"},{"key":"S026357472510177X_ref5","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2019.2896467"},{"key":"S026357472510177X_ref7","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2023.3310505"},{"key":"S026357472510177X_ref23","doi-asserted-by":"crossref","unstructured":"[23] Puang, E. Y. , Tee, K. P. and Jing, W. , \u201cKovis: Keypoint-based Visual Servoing with Zero-shot Sim-to-real Transfer for Robotics Manipulation,\u201d 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE (2020) pp. 7527\u20137533.","DOI":"10.1109\/IROS45743.2020.9341370"},{"key":"S026357472510177X_ref17","doi-asserted-by":"publisher","DOI":"10.1108\/IR-01-2021-0003"},{"key":"S026357472510177X_ref21","doi-asserted-by":"publisher","DOI":"10.1109\/MRA.2023.3282434"},{"key":"S026357472510177X_ref11","doi-asserted-by":"crossref","unstructured":"[11] Inoue, T. , De Magistris, G. , Munawar, A. , Yokoya, T. and Tachibana, R. , \u201cDeep Reinforcement Learning for High Precision Assembly Tasks,\u201d In: 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE (2017) pp. 819\u2013825.","DOI":"10.1109\/IROS.2017.8202244"},{"key":"S026357472510177X_ref25","first-page":"234","volume-title":"Medical Image Computing and Computer-assisted Intervention\u2013MICCAI 2015: 18th International Conference","author":"Ronneberger","year":"2015"},{"key":"S026357472510177X_ref4","unstructured":"[4] Bouchard, C. , Nesme, M. , Tournier, M. , Wang, B. , Faure, F. and Kry, P. , \u201c6d frictional contact for rigid bodies,\u201d In: Proceedings of the 41st Graphics Interface Conference, Canadian Information Processing Society(2015) pp. 105\u2013114."},{"key":"S026357472510177X_ref18","doi-asserted-by":"publisher","DOI":"10.1017\/S0263574723000425"},{"key":"S026357472510177X_ref41","doi-asserted-by":"crossref","unstructured":"[41] Zhao, W. , Queralta, J. P. and Westerlund, T. , \u201cSIm-to-real Transfer in Deep Reinforcement Learning for Robotics: A Survey,\u201d In: 2020 IEEE Symposium Series on Computational Intelligence (SSCI), IEEE (2020) pp. 737\u2013744.","DOI":"10.1109\/SSCI47803.2020.9308468"},{"key":"S026357472510177X_ref14","doi-asserted-by":"publisher","DOI":"10.1007\/s12541-014-0353-6"},{"key":"S026357472510177X_ref12","doi-asserted-by":"crossref","unstructured":"[12] James, S. , Wohlhart, P. , Kalakrishnan, M. , Kalashnikov, D. , Irpan, A. , Ibarz, J. , Levine, S. , Hadsell, R. and Bousmalis, K. , \u201cSim-to-real Via Sim-to-sim: Data-efficient Robotic Grasping Via Randomized-to-canonical Adaptation Networks,\u201d In: Proceedings Of The IEEE\/CVF Conference On Computer Vision And Pattern Recognition, IEEE (2019) pp. 12627\u201312637.","DOI":"10.1109\/CVPR.2019.01291"},{"key":"S026357472510177X_ref16","doi-asserted-by":"publisher","DOI":"10.1002\/aisy.202100095"},{"key":"S026357472510177X_ref19","doi-asserted-by":"crossref","unstructured":"[19] Luo, J. , Solowjow, E. , Wen, C. , Ojea, J. A. , Agogino, A. M. , Tamar, A. and Abbeel, P. , \u201cReinforcement Learning on Variable Impedance Controller for High-precision Robotic Assembly,\u201d In: 2019 International Conference on Robotics and Automation (ICRA), IEEE (2019) pp. 3080\u20133087.","DOI":"10.1109\/ICRA.2019.8793506"},{"key":"S026357472510177X_ref38","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2019.2940276"},{"key":"S026357472510177X_ref9","first-page":"4912","volume-title":"Proceedings of the 27th International Joint Conference on Artificial Intelligence (IJCAI)","author":"Zhang","year":"2018"},{"key":"S026357472510177X_ref27","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3076971"},{"key":"S026357472510177X_ref10","doi-asserted-by":"crossref","unstructured":"[10] Hou, Z. , Dong, H. , Zhang, K. , Gao, Q. , Chen, K. and Xu, J. , \u201cKnowledge-driven Deep Deterministic Policy Gradient for Robotic Multiple Peg-in-hole Assembly Tasks,\u201d In: 2018 IEEE International Conference on Robotics and Biomimetics (ROBIO) 2018, IEEE (2018) pp. 256\u2013261.","DOI":"10.1109\/ROBIO.2018.8665255"},{"key":"S026357472510177X_ref22","doi-asserted-by":"crossref","unstructured":"[22] Pashevich, A. , Strudel, R. , Kalevatykh, I. , Laptev, I. and Schmid, C. , \u201cLearning to Augment Synthetic Images for Sim2real Policy Transfer,\u201d In: 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE (2019) pp. 2651\u20132657.","DOI":"10.1109\/IROS40897.2019.8967622"},{"key":"S026357472510177X_ref1","doi-asserted-by":"publisher","DOI":"10.3390\/app10196923"},{"key":"S026357472510177X_ref35","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2018.2868859"},{"key":"S026357472510177X_ref33","doi-asserted-by":"publisher","DOI":"10.1017\/S0263574722000200"},{"key":"S026357472510177X_ref24","doi-asserted-by":"crossref","unstructured":"[24] Rao, K. , Harris, C. , Irpan, A. , Levine, S. , Ibarz, J. and Khansari, M. , \u201cRl-cyclegan: Reinforcement Learning Aware Simulation-to-real,\u201d In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE (2020) pp. 11157\u201311166.","DOI":"10.1109\/CVPR42600.2020.01117"},{"key":"S026357472510177X_ref15","first-page":"1","article-title":"A review of robot learning for manipulation: Challenges, representations, and algorithms","volume":"22","author":"Kroemer","year":"2021","journal-title":"J. Mach. Learn. Res."},{"key":"S026357472510177X_ref26","doi-asserted-by":"publisher","DOI":"10.1109\/TCDS.2023.3237734"},{"key":"S026357472510177X_ref36","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3187276"},{"key":"S026357472510177X_ref20","unstructured":"[20] Mnih, V. , Badia, A. P. , Mirza, M. , Graves, A. , Lillicrap, T. , Harley, T. , Silver, D. and Kavukcuoglu, K. ,\u201cAsynchronous Methods for Deep Reinforcement Learning,\u201d In: International Conference on Machine Learning, PMLR (2016) pp. 1928\u20131937."},{"key":"S026357472510177X_ref39","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2022.3184258"},{"key":"S026357472510177X_ref31","unstructured":"[31] Wang, Z. , Schaul, T. , Hessel, M. , Hasselt, H. , Lanctot, M. and Freitas, N. , \u201cDueling Network Architectures for Deep Reinforcement Learning,\u201d International Conference on Machine Learning, PMLR (2016) pp. 1995\u20132003."},{"key":"S026357472510177X_ref3","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.3011379"},{"key":"S026357472510177X_ref8","first-page":"1861","volume-title":"International Conference on Machine Learning","author":"Haarnoja","year":"2018"},{"key":"S026357472510177X_ref32","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3314762"},{"key":"S026357472510177X_ref13","doi-asserted-by":"crossref","unstructured":"[13] Jeong, R. , Aytar, Y. , Khosid, D. , Zhou, Y. , Kay, J. , Lampe, T. , Bousmalis, K. and Nori, F. , \u201cSelf-supervised Sim-to-real Adaptation for Visual Robotic Manipulation,\u201d In: 2020 IEEE International Conference on Robotics and Automation (ICRA), IEEE (2020) pp. 2718\u20132724.","DOI":"10.1109\/ICRA40945.2020.9197326"},{"key":"S026357472510177X_ref6","doi-asserted-by":"publisher","DOI":"10.1109\/JSEN.2021.3123638"},{"key":"S026357472510177X_ref29","doi-asserted-by":"crossref","unstructured":"[29] Tobin, J. , Fong, R. , Ray, A. , Schneider, J. , Zaremba, W. and Abbeel, P. , \u201cDomain Randomization for Transferring Deep Neural Networks From Simulation to the Real World,\u201d In: 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE (2017) pp. 23\u201330.","DOI":"10.1109\/IROS.2017.8202133"},{"key":"S026357472510177X_ref34","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3278208"},{"key":"S026357472510177X_ref37","doi-asserted-by":"publisher","DOI":"10.1109\/TIE.2020.3016271"},{"key":"S026357472510177X_ref30","doi-asserted-by":"crossref","unstructured":"[30] Villagomez, R. C. and Ordo\u00f1ez, J. , \u201cRobot Grasping Based on RGB Object And Grasp Detection Using Deep Learning,\u201d In: 2022, 8th International Conference on Mechatronics and Robotics Engineering (ICMRE), IEEE (2022) pp. 84\u201390.","DOI":"10.1109\/ICMRE54455.2022.9734075"}],"container-title":["Robotica"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S026357472510177X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T08:25:02Z","timestamp":1757579102000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S026357472510177X\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,18]]},"references-count":41,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["S026357472510177X"],"URL":"https:\/\/doi.org\/10.1017\/s026357472510177x","relation":{},"ISSN":["0263-5747","1469-8668"],"issn-type":[{"value":"0263-5747","type":"print"},{"value":"1469-8668","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,18]]}}}