{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,11]],"date-time":"2025-12-11T09:25:15Z","timestamp":1765445115155,"version":"3.46.0"},"reference-count":42,"publisher":"Emerald","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,11,25]]},"abstract":"<jats:sec>\n                    <jats:title>Purpose<\/jats:title>\n                    <jats:p>The purpose of this study is to address the challenge of object manipulation in scenarios where the target is not explicitly defined, requiring robots to engage in efficient planning to determine the sequence of actions for picking, placing and positioning objects. The aim is to develop a multistep skill learning method that integrates perception with a set of primitive actions, including a novel action of orienting, to enable robots to perform complex tasks that require multistep planning and interaction with various objects in cluttered and unstructured environments.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Design\/methodology\/approach<\/jats:title>\n                    <jats:p>To achieve the purpose, the authors propose a pipeline that decomposes the object manipulation task into three independent stages, each trained end-to-end with raw visual inputs using off-policy reinforcement learning algorithms. The Q-learning algorithm is used to simultaneously train three fully convolutional neural networks for each primitive action \u2013 grasping, pushing, placing and orienting \u2013 from scratch. The framework is designed to be modular, allowing for easy extension to multistep manipulation tasks.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Findings<\/jats:title>\n                    <jats:p>The findings demonstrate that robots can learn complex behaviors through both simulated and real-world experiments. In simulation, the robot achieved an efficient block-stacking success rate of up to 98% during testing. When transferring the model to a real universal robots UR3 (UR3) robot using effective domain randomization, the robot achieved a 100% completion rate with convex objects and a 92% completion rate with various objects not seen during training.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Originality\/value<\/jats:title>\n                    <jats:p>We develop a novel multistep skill learning method that integrates perception with multiple primitive actions, including a new action of orienting, and the use of off-policy reinforcement learning algorithms for end-to-end training. The modular design of the framework allows for easy extension to more complex manipulation tasks, and the encouraging results in both simulated and real-world experiments demonstrate significant improvements over current long-term planning methods.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1108\/ir-12-2024-0534","type":"journal-article","created":{"date-parts":[[2025,3,28]],"date-time":"2025-03-28T01:07:32Z","timestamp":1743124052000},"page":"853-865","source":"Crossref","is-referenced-by-count":0,"title":["Self-supervised reinforcement learning for multi-step object manipulation skills"],"prefix":"10.1108","volume":"52","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-5048-2572","authenticated-orcid":true,"given":"Jiaqi","family":"Wang","sequence":"first","affiliation":[{"name":"South China University of Technology , Guangzhou,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0008-5302","authenticated-orcid":true,"given":"Chuxin","family":"Chen","sequence":"additional","affiliation":[{"name":"South China University of Technology , Guangzhou,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingwei","family":"Liu","sequence":"additional","affiliation":[{"name":"South China University of Technology , Guangzhou,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9425-843X","authenticated-orcid":true,"given":"Guanglong","family":"Du","sequence":"additional","affiliation":[{"name":"South China University of Technology Department of CS, , Guangzhou,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9506-9005","authenticated-orcid":true,"given":"Xiaojun","family":"Zhu","sequence":"additional","affiliation":[{"name":"Jianghuai Advanced Technology Center , Hefei,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Quanlong","family":"Guan","sequence":"additional","affiliation":[{"name":"Jinan University , Guangzhou,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaojian","family":"Qiu","sequence":"additional","affiliation":[{"name":"Institute for Military-Civilian Integration of Jiangxi Province , Nanchang,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","published-online":{"date-parts":[[2025,3,31]]},"reference":[{"key":"2025121104230384600_ref001","first-page":"367","article-title":"Current research trends in robot grasping and bin picking[C]","author":"Alonso","year":"2019"},{"key":"2025121104230384600_ref002","article-title":"Never give up: learning directed exploration strategies[J]","author":"Badia","year":"2020","journal-title":"arXiv Preprint arXiv:2002.06038."},{"key":"2025121104230384600_ref003","first-page":"612","article-title":"Robot learning of shifting objects for grasping in cluttered environments[C]","author":"Berscheid","year":"2019"},{"key":"2025121104230384600_ref004","first-page":"4243","article-title":"Using simulation and domain adaptation to improve efficiency of deep robotic grasping[C]\/\/2018","author":"Bousmalis","year":"2018"},{"key":"2025121104230384600_ref005","first-page":"4895","article-title":"Efficient heatmap-guided 6-DoF grasp detection in cluttered scenes[J]. IEEE robotics and automation letters","author":"Chen","year":"2023"},{"key":"2025121104230384600_ref006","article-title":"Reducing the barrier to entry of complex robotic software: a moveit! case study[J]","author":"Coleman","year":"2014","journal-title":"arXiv Preprint arXiv:1404.3785."},{"key":"2025121104230384600_ref007","first-page":"1559","article-title":"Neural networks for incremental dimensionality reduced reinforcement learning[C]","author":"Curran","year":"2017"},{"issue":"12","key":"2025121104230384600_ref008","doi-asserted-by":"crossref","first-page":"3846","DOI":"10.1017\/S0263574723001285","article-title":"A review of robotic grasp detection technology[J]","volume":"41","author":"Dong","year":"2023","journal-title":"Robotica"},{"key":"2025121104230384600_ref009","first-page":"405","article-title":"Learning to singulate objects using a push proposal network[C]","author":"Eitel","year":"2020"},{"key":"2025121104230384600_ref010","first-page":"7433","article-title":"Pick and place without geometric object models[C]","author":"Gualtieri","year":"2018"},{"issue":"7","key":"2025121104230384600_ref011","doi-asserted-by":"crossref","first-page":"3762","DOI":"10.3390\/s23073762","article-title":"A survey on deep reinforcement learning algorithms for robotic manipulation[J]","volume":"23","author":"Han","year":"2023","journal-title":"Sensors"},{"key":"2025121104230384600_ref012","first-page":"770","article-title":"Deep residual learning for image recognition[C]","author":"He","year":"2016"},{"issue":"4","key":"2025121104230384600_ref013","doi-asserted-by":"crossref","first-page":"6724","DOI":"10.1109\/LRA.2020.3015448","article-title":"Good robot!\u201d: efficient reinforcement learning for multi-step visual tasks with sim to real transfer[J]","volume":"5","author":"Hundt","year":"2020","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2025121104230384600_ref014","first-page":"334","article-title":"Transferring end-to-end visuomotor control from simulation to real world for a multi-stage task[C]","author":"James","year":"2017"},{"key":"2025121104230384600_ref015","first-page":"651","article-title":"Scalable deep reinforcement learning for vision-based robotic manipulation","author":"Kalashnikov","year":"2018"},{"key":"2025121104230384600_ref016","article-title":"A method for stochastic optimization[J]","author":"Kingma","year":"2014","journal-title":"arXiv Preprint arXiv:1412.6980."},{"issue":"30","key":"2025121104230384600_ref017","first-page":"1","article-title":"A review of robot learning for manipulation: challenges, representations, and algorithms[J]","volume":"22","author":"Kroemer","year":"2021","journal-title":"Journal of Machine Learning Research"},{"key":"2025121104230384600_ref018","first-page":"995","article-title":"RRT-connect: an efficient approach to single-query path planning[C]","author":"Kuffner","year":"2000"},{"issue":"1","key":"2025121104230384600_ref019","doi-asserted-by":"crossref","first-page":"534","DOI":"10.1109\/LRA.2021.3129833","article-title":"Learning robotic manipulation tasks via task progress based Gaussian reward and loss adjusted exploration[J]","volume":"7","author":"Kumra","year":"2021","journal-title":"IEEE Robotics and Automation Letters"},{"issue":"2","key":"2025121104230384600_ref020","doi-asserted-by":"crossref","DOI":"10.1063\/1.5006570","article-title":"A novel algorithm for fast grasping of unknown objects using C-shape configuration[J]","volume":"8","author":"Lei","year":"2018","journal-title":"AIP Advances"},{"issue":"5","key":"2025121104230384600_ref021","doi-asserted-by":"crossref","first-page":"773","DOI":"10.1007\/s11370-021-00398-z","article-title":"A survey on deep learning and deep reinforcement learning in robotics with a tutorial on deep reinforcement learning[J]","volume":"14","author":"Morales","year":"2021","journal-title":"Intelligent Service Robotics"},{"key":"2025121104230384600_ref022","first-page":"807","article-title":"Rectified linear units improve restricted Boltzmann machines[C]","author":"Nair","year":"2010"},{"key":"2025121104230384600_ref023","first-page":"1321","article-title":"V-REP: a versatile and scalable robot simulation framework[C]","author":"Rohmer","year":"2013"},{"issue":"3","key":"2025121104230384600_ref024","doi-asserted-by":"crossref","first-page":"1711","DOI":"10.1109\/LRA.2018.2801939","article-title":"Nonprehensile dynamic manipulation: a survey[J]","volume":"3","author":"Ruggiero","year":"2018","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2025121104230384600_ref025","first-page":"31","article-title":"How does batch normalization help optimization?[J]","author":"Santurkar","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2025121104230384600_ref026","first-page":"6225","article-title":"Split deep q-learning for robust object singulation[C]","author":"Sarantopoulos","year":"2020"},{"key":"2025121104230384600_ref027","first-page":"1513","article-title":"Push-to-see: learning non-prehensile manipulation to enhance instance segmentation via deep q-learning[C]","author":"Serhan","year":"2022"},{"key":"2025121104230384600_ref028","article-title":"A comparative study on state-action spaces for learning viewpoint selection and manipulation with diffusion policy[J]","author":"Sun","year":"2024","journal-title":"arXiv Preprint arXiv:2409.14615."},{"issue":"1","key":"2025121104230384600_ref029","first-page":"9","volume":"1","author":"Sutton","year":"1998","journal-title":"Reinforcement Learning: An Introduction[J]"},{"key":"2025121104230384600_ref030","first-page":"291","article-title":"Learning a visuomotor controller for real world robotic grasping using simulated depth images[C]","author":"Viereck","year":"2017"},{"key":"2025121104230384600_ref031","first-page":"606","article-title":"Towards privacy-preserving visual recognition via adversarial training: a pilot study[C]\/\/Proceedings of the European Conference on Computer Vision (ECCV)","author":"Wu","year":"2018"},{"key":"2025121104230384600_ref032","doi-asserted-by":"crossref","first-page":"1038658","DOI":"10.3389\/frobt.2023.1038658","article-title":"Learning-based robotic grasping: a review[J]","volume":"10","author":"Xie","year":"2023","journal-title":"Frontiers in Robotics and AI"},{"issue":"4","key":"2025121104230384600_ref033","doi-asserted-by":"crossref","first-page":"1730","DOI":"10.1109\/TASE.2020.3017022","article-title":"Robotic grasping of unknown objects using novel multilevel convolutional neural networks: from parallel gripper to dexterous hand[J]","volume":"18","author":"Yu","year":"2020","journal-title":"IEEE Transactions on Automation Science and Engineering"},{"key":"2025121104230384600_ref034","first-page":"4238","article-title":"Learning synergies between pushing and grasping with self-supervised deep reinforcement learning[C]","author":"Zeng","year":"2018"},{"key":"2025121104230384600_ref035","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1016\/j.patrec.2020.09.014","article-title":"Visual manipulation relationship recognition in object-stacking scenes[J]","volume":"140","author":"Zhang","year":"2020","journal-title":"Pattern Recognition Letters"},{"issue":"10\/11","key":"2025121104230384600_ref036","doi-asserted-by":"crossref","first-page":"1229","DOI":"10.1177\/0278364919870227","article-title":"Adversarial discriminative sim-to-real transfer of visuo-motor policies[J]","volume":"38","author":"Zhang","year":"2019","journal-title":"The International Journal of Robotics Research"},{"issue":"2","key":"2025121104230384600_ref037","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1109\/MIM.2022.9756392","article-title":"Deep learning-based robot vision: high-end tools for smart manufacturing[J]","volume":"25","author":"Zhang","year":"2022","journal-title":"IEEE Instrumentation & Measurement Magazine"},{"key":"2025121104230384600_ref038","doi-asserted-by":"crossref","first-page":"102601","DOI":"10.1016\/j.rcim.2023.102601","article-title":"Digital twin-enabled grasp outcomes assessment for unknown objects using visual-tactile fusion perception[J]","volume":"84","author":"Zhang","year":"2023","journal-title":"Robotics and Computer-Integrated Manufacturing"},{"key":"2025121104230384600_ref039","first-page":"737","article-title":"Sim-to-real transfer in deep reinforcement learning for robotics: a survey[C]","author":"Zhao","year":"2020"},{"key":"2025121104230384600_ref040","first-page":"7223","article-title":"Fully convolutional grasp detection network with oriented anchor box[C]\/\/2018","author":"Zhou","year":"2018"},{"key":"2025121104230384600_ref041","doi-asserted-by":"crossref","first-page":"102541","DOI":"10.1016\/j.rcim.2023.102541","article-title":"Instance segmentation based 6D pose estimation of industrial objects using point clouds for robotic bin-picking[J]","volume":"82","author":"Zhuang","year":"2023","journal-title":"Robotics and Computer-Integrated Manufacturing"},{"key":"2025121104230384600_ref042","first-page":"477","article-title":"Learning 6-DoF grasping and pick-place using attention focus[C]","author":"Gualtieri","year":"2018"}],"container-title":["Industrial Robot: the international journal of robotics research and application"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/IR-12-2024-0534\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/ir\/article-pdf\/52\/6\/853\/10834348\/ir-12-2024-0534en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ir\/article-pdf\/52\/6\/853\/10834348\/ir-12-2024-0534en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,11]],"date-time":"2025-12-11T09:23:15Z","timestamp":1765444995000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ir\/article\/52\/6\/853\/1257358\/Self-supervised-reinforcement-learning-for-multi"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,31]]},"references-count":42,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,11,25]]}},"URL":"https:\/\/doi.org\/10.1108\/ir-12-2024-0534","relation":{},"ISSN":["0143-991X","1758-5791"],"issn-type":[{"type":"print","value":"0143-991X"},{"type":"electronic","value":"1758-5791"}],"subject":[],"published":{"date-parts":[[2025,3,31]]}}}