{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T16:20:49Z","timestamp":1776183649146,"version":"3.50.1"},"reference-count":35,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,7,9]],"date-time":"2024-07-09T00:00:00Z","timestamp":1720483200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,7,9]],"date-time":"2024-07-09T00:00:00Z","timestamp":1720483200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Intell Robot Syst"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>During the movement of a robotic arm, collisions can easily occur if the arm directly grasps at multiple tightly stacked objects, thereby leading to grasp failures or machine damage. Grasp success can be improved through the rearrangement or movement of objects to clear space for grasping. This paper presents a high-performance deep Q-learning framework that can help robotic arms to learn synchronized push and grasp tasks. In this framework, a grasp quality network is used for precisely identifying stable grasp positions on objects to expedite model convergence and solve the problem of sparse rewards caused during training because of grasp failures. Furthermore, a novel reward function is proposed for effectively evaluating whether a pushing action is effective. The proposed framework achieved grasp success rates of 92% and 89% in simulations and real-world experiments, respectively. Furthermore, only 200 training steps were required to achieve a grasp success rate of 80%, which indicates the suitability of the proposed framework for rapid deployment in industrial settings.<\/jats:p>","DOI":"10.1007\/s10846-024-02127-x","type":"journal-article","created":{"date-parts":[[2024,7,9]],"date-time":"2024-07-09T12:01:35Z","timestamp":1720526495000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Integration of Deep Q-Learning with a Grasp Quality Network for Robot Grasping in Cluttered Environments"],"prefix":"10.1007","volume":"110","author":[{"given":"Chih-Yung","family":"Huang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu-Hsiang","family":"Shao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,7,9]]},"reference":[{"issue":"10","key":"2127_CR1","doi-asserted-by":"publisher","first-page":"4873","DOI":"10.1109\/TCYB.2020.2998837","volume":"51","author":"L Kong","year":"2021","unstructured":"Kong, L., He, W., Yang, W., Li, Q., Kaynak, O.: Fuzzy approximation-based finite-time control for a robot with actuator saturation under time-varying constraints of work space. IEEE Trans. Cybern. 51(10), 4873\u20134884 (2021). https:\/\/doi.org\/10.1109\/TCYB.2020.2998837","journal-title":"IEEE Trans. Cybern."},{"issue":"8","key":"2127_CR2","doi-asserted-by":"publisher","first-page":"3052","DOI":"10.1109\/TCYB.2018.2838573","volume":"49","author":"L Kong","year":"2019","unstructured":"Kong, L., He, W., Yang, C., Li, Z., Sun, C.: adaptive fuzzy control for coordinated multiple robots with constraint using impedance learning. IEEE Trans. Cybern. 49(8), 3052\u20133063 (2019). https:\/\/doi.org\/10.1109\/TCYB.2018.2838573","journal-title":"IEEE Trans. Cybern."},{"issue":"3","key":"2127_CR3","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1109\/TSMC.2019.2901277","volume":"51","author":"L Kong","year":"2021","unstructured":"Kong, L., He, W., Dong, Y., Cheng, L., Yang, C., Li, Z.: Asymmetric bounded neural control for an uncertain robot by state feedback and output feedback. IEEE Trans. Syst. Man, Cybern. Syst. 51(3), 1735\u20131746 (2021). https:\/\/doi.org\/10.1109\/TSMC.2019.2901277","journal-title":"IEEE Trans. Syst. Man, Cybern. Syst."},{"key":"2127_CR4","doi-asserted-by":"publisher","unstructured":"Miller, A.T., Knoop, S., Christensen, H.I., Allen, P.K.: Automatic grasp planning using shape primitives. In: 2003 IEEE International Conference on Robotics and Automation (Cat. No.03CH37422), vol. 2, pp. 1824\u20131829 (2003). https:\/\/doi.org\/10.1109\/ROBOT.2003.1241860","DOI":"10.1109\/ROBOT.2003.1241860"},{"issue":"2","key":"2127_CR5","doi-asserted-by":"publisher","first-page":"289","DOI":"10.1109\/TRO.2013.2289018","volume":"30","author":"J Bohg","year":"2014","unstructured":"Bohg, J., Morales, A., Asfour, T., Kragic, D.: Data-driven grasp synthesis\u2014a survey. IEEE Trans. Rob. 30(2), 289\u2013309 (2014). https:\/\/doi.org\/10.1109\/TRO.2013.2289018","journal-title":"IEEE Trans. Rob."},{"key":"2127_CR6","doi-asserted-by":"publisher","unstructured":"Tian, H., Song, K., Li, S., Ma, S., Xu, J., Yan, Y.: Data-driven robotic visual grasping detection for unknown objects: a problem-oriented review. Exp. Syst. Appl. 211, 118624 (2023). https:\/\/doi.org\/10.1016\/j.eswa.2022.118624","DOI":"10.1016\/j.eswa.2022.118624"},{"key":"2127_CR7","doi-asserted-by":"publisher","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., Girshick, R.: Mask R-CNN. In: 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2980-2988 (2017). https:\/\/doi.org\/10.1109\/ICCV.2017.322","DOI":"10.1109\/ICCV.2017.322"},{"key":"2127_CR8","doi-asserted-by":"publisher","unstructured":"Cao, Z., Simon, T., Wei, S.E., Sheikh, Y.: Realtime multi-person 2D pose estimation using part affinity fields. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1302-1310 (2017). https:\/\/doi.org\/10.1109\/CVPR.2017.143","DOI":"10.1109\/CVPR.2017.143"},{"key":"2127_CR9","doi-asserted-by":"publisher","unstructured":"Mnih, V. et al.: Human-level control through deep reinforcement learning. Nature 518(7540), 529\u2013533 (2015). https:\/\/doi.org\/10.1038\/nature14236","DOI":"10.1038\/nature14236"},{"key":"2127_CR10","doi-asserted-by":"publisher","unstructured":"Zeng, A., Song, S., Welker, S., Lee, J., Rodriguez, A., Funkhouser, T.: Learning synergies between pushing and grasping with self-supervised deep reinforcement learning. In: 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 4238-4245 (2018). https:\/\/doi.org\/10.1109\/IROS.2018.8593986","DOI":"10.1109\/IROS.2018.8593986"},{"key":"2127_CR11","doi-asserted-by":"publisher","unstructured":"Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3431-3440 (2015). https:\/\/doi.org\/10.1109\/CVPR.2015.7298965","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"2127_CR12","doi-asserted-by":"publisher","unstructured":"Berscheid, L., Mei\u00dfner, P., Kr\u00f6ger, T.: Robot Learning of Shifting Objects for Grasping in Cluttered Environments. In: 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 612-618 (2019). https:\/\/doi.org\/10.1109\/IROS40897.2019.8968042","DOI":"10.1109\/IROS40897.2019.8968042"},{"issue":"1","key":"2127_CR13","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1109\/JAS.2021.1004255","volume":"9","author":"Y Yang","year":"2022","unstructured":"Yang, Y., Ni, Z., Gao, M., Zhang, J., Tao, D.: Collaborative pushing and grasping of tightly stacked objects via deep reinforcement learning. IEEE\/CAA J. Autom. Sin. 9(1), 135\u2013145 (2022). https:\/\/doi.org\/10.1109\/JAS.2021.1004255","journal-title":"IEEE\/CAA J. Autom. Sin."},{"issue":"4","key":"2127_CR14","doi-asserted-by":"publisher","first-page":"6337","DOI":"10.1109\/LRA.2021.3092640","volume":"6","author":"K Xu","year":"2021","unstructured":"Xu, K., Yu, H., Lai, Q., Wang, Y., Xiong, R.: Efficient learning of goal-oriented push-grasping synergy in clutter. IEEE Robot. Autom. Lett. 6(4), 6337\u20136344 (2021). https:\/\/doi.org\/10.1109\/LRA.2021.3092640","journal-title":"IEEE Robot. Autom. Lett."},{"key":"2127_CR15","doi-asserted-by":"publisher","unstructured":"Mahler, J. et al.:Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics (2017). https:\/\/doi.org\/10.48550\/arXiv.1703.09312","DOI":"10.48550\/arXiv.1703.09312"},{"key":"2127_CR16","doi-asserted-by":"publisher","unstructured":"Sahbani, A., El-Khoury, S., Bidaud, P.: An overview of 3D object grasp synthesis algorithms. Robot. Auton. Syst. 60(3), 326\u2013336 (2012). https:\/\/doi.org\/10.1016\/j.robot.2011.07.016","DOI":"10.1016\/j.robot.2011.07.016"},{"issue":"4","key":"2127_CR17","doi-asserted-by":"publisher","first-page":"3355","DOI":"10.1109\/LRA.2018.2852777","volume":"3","author":"FJ Chu","year":"2018","unstructured":"Chu, F.J., Xu, R., Vela, P.A.: Real-world multiobject, multigrasp detection. IEEE Rob. Autom. Lett. 3(4), 3355\u20133362 (2018). https:\/\/doi.org\/10.1109\/LRA.2018.2852777","journal-title":"IEEE Rob. Autom. Lett."},{"key":"2127_CR18","doi-asserted-by":"publisher","unstructured":"Morrison, D., Corke, P., Leitner, J.: Closing the loop for robotic grasping: A real-time, generative grasp synthesis approach (2018). https:\/\/doi.org\/10.48550\/arXiv.1804.05172","DOI":"10.48550\/arXiv.1804.05172"},{"key":"2127_CR19","doi-asserted-by":"publisher","unstructured":"Sundermeyer, M., Mousavian, A., Triebel, R., Fox, R.: Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes. In: 2021 IEEE International Conference on Robotics and Automation (ICRA), pp. 13438-13444 (2021). https:\/\/doi.org\/10.1109\/ICRA48506.2021.9561877","DOI":"10.1109\/ICRA48506.2021.9561877"},{"key":"2127_CR20","doi-asserted-by":"publisher","unstructured":"Ghadirzadeh, A., Maki, A., Kragic, D., Bj\u00f6rkman, M.: Deep predictive policy training using reinforcement learning. In: 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 2351-2358 (2017). https:\/\/doi.org\/10.1109\/IROS.2017.8206046","DOI":"10.1109\/IROS.2017.8206046"},{"key":"2127_CR21","doi-asserted-by":"publisher","unstructured":"Quillen, D., Jang, E., Nachum, O., Finn, C., Ibarz, J., Levine, S.: Deep Reinforcement Learning for Vision-Based Robotic Grasping: A Simulated Comparative Evaluation of Off-Policy Methods. In: 2018 IEEE International Conference on Robotics and Automation (ICRA), pp. 6284-6291 (2018). https:\/\/doi.org\/10.1109\/ICRA.2018.8461039","DOI":"10.1109\/ICRA.2018.8461039"},{"key":"2127_CR22","doi-asserted-by":"publisher","unstructured":"Yen-Chen, L., Zeng, A., Song, S., Isola, P., Lin, T.Y.: Learning to See before Learning to Act: Visual Pre-training for Manipulation. In: 2020 IEEE International Conference on Robotics and Automation (ICRA), pp. 7286-7293 (2020). https:\/\/doi.org\/10.1109\/ICRA40945.2020.9197331","DOI":"10.1109\/ICRA40945.2020.9197331"},{"key":"2127_CR23","doi-asserted-by":"publisher","unstructured":"Kalashnikov, D. et al.: Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation (2018). https:\/\/doi.org\/10.48550\/arXiv.1806.10293","DOI":"10.48550\/arXiv.1806.10293"},{"issue":"2","key":"2127_CR24","doi-asserted-by":"publisher","first-page":"2232","DOI":"10.1109\/LRA.2020.2970622","volume":"5","author":"Y Yang","year":"2020","unstructured":"Yang, Y., Liang, H., Choi, C.: A deep learning approach to grasping the invisible. IEEE Robot. Autom. Lett. 5(2), 2232\u20132239 (2020). https:\/\/doi.org\/10.1109\/LRA.2020.2970622","journal-title":"IEEE Robot. Autom. Lett."},{"issue":"4","key":"2127_CR25","doi-asserted-by":"publisher","first-page":"11966","DOI":"10.1109\/LRA.2022.3204822","volume":"7","author":"E Li","year":"2022","unstructured":"Li, E., Feng, H., Zhang, S., Fu, Y.: Learning target-oriented push-grasping synergy in clutter with action space decoupling. IEEE Robot. Autom. Lett. 7(4), 11966\u201311973 (2022). https:\/\/doi.org\/10.1109\/LRA.2022.3204822","journal-title":"IEEE Robot. Autom. Lett."},{"key":"2127_CR26","doi-asserted-by":"publisher","unstructured":"Florence, P., et al.: Implicit behavioral cloning. In: Conference on Robot Learning, pp. 158\u2013168: PMLR (2022). https:\/\/doi.org\/10.48550\/arXiv.2109.00137","DOI":"10.48550\/arXiv.2109.00137"},{"key":"2127_CR27","doi-asserted-by":"publisher","unstructured":"Zeng, A., et al.: Transporter networks: Rearranging the visual world for robotic manipulation. In: Conference on Robot Learning, pp. 726\u2013747 : PMLR (2021). https:\/\/doi.org\/10.48550\/arXiv.2010.14406","DOI":"10.48550\/arXiv.2010.14406"},{"key":"2127_CR28","unstructured":"Nair, V., Hinton, G.E.: Rectified linear units improve restricted boltzmann machines. Presented at the Proceedings of the 27th International Conference on International Conference on Machine Learning, Haifa, Israel (2010)"},{"key":"2127_CR29","doi-asserted-by":"publisher","unstructured":"Huang, B., Han, S.D., Boularias, A., Yu, J.: DIPN: Deep Interaction Prediction Network with Application to Clutter Removal. In: 2021 IEEE International Conference on Robotics and Automation (ICRA), pp. 4694-4701 (2021). https:\/\/doi.org\/10.1109\/ICRA48506.2021.9561073","DOI":"10.1109\/ICRA48506.2021.9561073"},{"key":"2127_CR30","doi-asserted-by":"publisher","unstructured":"He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770\u2013778 (2016). https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"2127_CR31","doi-asserted-by":"publisher","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2117\u20132125 (2017). https:\/\/doi.org\/10.1109\/CVPR.2017.106","DOI":"10.1109\/CVPR.2017.106"},{"key":"2127_CR32","doi-asserted-by":"publisher","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q.: Densely connected convolutional networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700\u20134708 (2017). https:\/\/doi.org\/10.1109\/CVPR.2017.243","DOI":"10.1109\/CVPR.2017.243"},{"key":"2127_CR33","doi-asserted-by":"publisher","unstructured":"Schaul, T., Quan, J., Antonoglou, I., Silver, D.: Prioritized experience replay. (2015). https:\/\/doi.org\/10.48550\/arXiv.1511.05952","DOI":"10.48550\/arXiv.1511.05952"},{"key":"2127_CR34","doi-asserted-by":"publisher","unstructured":"Rohmer, E., Singh, S. P. N., Freese, M.: V-REP: A versatile and scalable robot simulation framework. In: 2013 IEEE\/RSJ International Conference on Intelligent Robots and Systems, pp. 1321-1326 (2013). https:\/\/doi.org\/10.1109\/IROS.2013.6696520","DOI":"10.1109\/IROS.2013.6696520"},{"issue":"4","key":"2127_CR35","doi-asserted-by":"publisher","first-page":"6724","DOI":"10.1109\/LRA.2020.3015448","volume":"5","author":"A Hundt","year":"2020","unstructured":"Hundt, A., et al.: \u201cGood Robot!\u201d: efficient reinforcement learning for multi-step visual tasks with sim to real transfer. IEEE Robot. Autom. Lett. 5(4), 6724\u20136731 (2020). https:\/\/doi.org\/10.1109\/LRA.2020.3015448","journal-title":"IEEE Robot. Autom. Lett."}],"container-title":["Journal of Intelligent &amp; Robotic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10846-024-02127-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10846-024-02127-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10846-024-02127-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,26]],"date-time":"2024-09-26T13:09:10Z","timestamp":1727356150000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10846-024-02127-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,9]]},"references-count":35,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["2127"],"URL":"https:\/\/doi.org\/10.1007\/s10846-024-02127-x","relation":{},"ISSN":["1573-0409"],"issn-type":[{"value":"1573-0409","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,9]]},"assertion":[{"value":"6 December 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 June 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 July 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics Approval"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to Participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to Publish"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing Interests"}}],"article-number":"97"}}